AI News

Qwen3.8-27B vs Qwen3.6-27B vs Gemma 4 31B: Which Is Best for a 24GB GPU?

Short verdict

For most people buying or repurposing a 24GB GPU for serious local work, Qwen3.8-27B is the best default of these three. On one RTX 4090, its Q4_K_M file matched Qwen3.6-27B’s roughly 49-token/s decode speed and memory curve, yet completed all 12 coding tasks on the first declared seed, passed 39/40 reasoning cases across three seeds, and answered 23/24 document questions. That is a substantial practical upgrade over Qwen3.6 without a meaningful throughput penalty.

Gemma 4 31B-it is the credible alternative. It posted the cleanest structured-tool result—90/90 single calls and 30/30 multi-step runs—and the best vision count at 19/20. It also edged the blinded writing rubric by one point. The costs were higher VRAM use, about 8.3% slower synthetic decode, a slower warm time-to-first-token, and weaker coding-agent completion.

Qwen3.6-27B still makes sense when an existing integration is stable or tightly tuned to its template. For a new deployment, however, it is difficult to recommend over Qwen3.8.

These conclusions apply to the exact pinned GGUFs and runtime below. The highest stable context we tested was 64K—not the models’ advertised 262K—and “fits” does not mean “comfortable” once projectors, speculative decoding, a desktop display, or concurrency consume the remaining VRAM. See Qwen3.8-27B hardware requirements for other hardware tiers.

The one-minute chooser

There is no single best answer independent of workload. The equal-Q4 lane asks how the models behave when all three use the same quantization family, full GPU offload and F16 K/V cache. The best-fit lane asks what each can sustain inside the same 24GB ceiling. Those are different questions: Gemma’s 64K profile, for example, failed with F16 K/V but passed after moving its K/V cache to Q8_0.

If your priority is… Pick Why in this test Important caveat
Coding agents and repository-style edits Qwen3.8-27B 12/12 pass@1; 35/36 seeded runs Small synthetic Python tasks, not a full production repository benchmark
Reasoning under a fixed token budget Qwen3.8-27B 39/40 pass@3; 117/120 runs A larger budget can change results, especially for Qwen3.6
Private document QA Qwen3.8-27B 23/24, narrowly ahead of Gemma’s 22/24 Tested at 8K and 32K with synthetic evidence packs
Strict, declared tool calls Gemma 4 31B-it 90/90 single and 30/30 multi-step runs Its iterative coding-tool loop was much less reliable
Vision-heavy local work Gemma 4 31B-it 19/20, versus 18/20 for Qwen3.8 Small synthetic image set; no video or audio verdict
Writing and copyediting Tie in practice Blinded scores were 94/96, 93/96 and 93/96 One evaluator and 12 samples per model
64K context with the most headroom Qwen3.8 or Qwen3.6 4.20 GiB free with F16 K/V Qwen3.6’s task reliability was much lower in key lanes
Minimum migration work Qwen3.6-27B Keep an already validated stack Not the strongest choice for a fresh install

If your real constraint is a different card, system RAM, or CPU-offload tolerance, use the broader local AI compatibility guide rather than transferring these RTX 4090 numbers directly.

Why these three models—and why not every 24GB model?

The comparison is deliberately narrow. Qwen3.8-27B is the focal release. Qwen3.6-27B is its direct dense, multimodal predecessor in the same size class. Gemma 4 31B-it is the closest non-Qwen dense multimodal peer that can plausibly live fully on a 24GB card at Q4.

That design isolates a useful purchasing and upgrade question. Qwen3.8 and Qwen3.6 share the same broad 64-layer hybrid pattern, token vocabulary and vision-encoder depth, so their comparison is unusually clean. Gemma supplies a different architecture and training lineage without moving into server-class memory requirements.

We excluded speed-first mixture-of-experts models such as Qwen3.6-35B-A3B and Gemma 4 26B-A4B because active-parameter throughput changes the question. We also excluded older dense checkpoints and larger models that require CPU offload. They may be worthwhile; they simply belong in a separate speed-first or multi-tier comparison. Kingy.ai’s broader open-weight model comparison covers a wider field.

We use open-weight as the precise umbrella term. All three pinned model cards list Apache 2.0, but “open source” can also imply disclosed training data, complete recipes and reproducibility beyond downloadable weights. That distinction matters when evaluating openness as well as local usability.

Model identity, specifications and exact files

Identity mistakes can invalidate a comparison. The Gemma checkpoint here is the instruction-tuned google/gemma-4-31B-it, not the base google/gemma-4-31B. The Qwen repositories are the post-trained 27B checkpoints, not similarly named APIs or smaller variants. The full Qwen3.8-27B specifications and launch analysis provides broader launch context; this page stays focused on the same-GPU decision.

Derived from pinned configuration, both Qwen models have 64 text layers: 48 linear-attention layers and 16 full-attention layers, with full attention every fourth layer. Their vision encoders have 27 layers. Gemma has 60 text layers—50 sliding-attention and 10 full-attention layers—and a 27-layer vision encoder. All three configurations declare 262,144 maximum positions. That declaration is an architectural limit, not a promise that a Q4 build can allocate the necessary cache on a 24GB card.

The runtime counted 27.32 billion parameters for Qwen3.8, 26.90 billion for Qwen3.6 and 30.70 billion for Gemma. Exact artifacts are below. Sizes are binary GiB; SHA-256 values were checked before testing.

Model / artifact Pinned revision File and size SHA-256
Qwen3.8-27B source 1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0 Post-trained model record Config snapshot retained in evidence bundle
Qwen3.8 Q4 f1bfb127c64f7072bdd2cad55f258b9c8b2910fe Qwen3.8-27B-Q4_K_M.gguf, 15.932 GiB 7e78da5d7e3ae28d178121f58646953305f3e5bd3cb46f4a75584e8b6c6fe169
Qwen3.6-27B source 6a9e13bd6fc8f0983b9b99948120bc37f49c13e9 Post-trained model record Config snapshot retained in evidence bundle
Qwen3.6 Q4 82d411acf4a06cfb8d9b073a5211bf410bfc29bf Qwen3.6-27B-Q4_K_M.gguf, 15.662 GiB 5ed60d0af4650a854b1755bd392f9aef4872643dc25a254bc68043fa638392a0
Gemma 4 source 842da3794eaa0b77d5f08bae87a17459d91ff475 Instruction-tuned model record Config snapshot retained in evidence bundle
Gemma 4 Q4 c1ac76e99d5513b141e8adde7288b85c3f9c32ec gemma-4-31B-it-Q4_K_M.gguf, 17.065 GiB 38bd64c852c4b460434cc7162fa9bdcf242faf86502581a754cb72956bb17f84

Vision was optional, not silently baked into text tests. The Qwen projectors were 0.864 GiB each (Qwen3.8 projector, Qwen3.6 projector); the Gemma projector was 1.117 GiB. Those are file sizes, not claims that peak VRAM rises by exactly the same amount.

How Kingy.ai tested the models

Kingy.ai measured all three on one physical NVIDIA GeForce RTX 4090 reporting 24,564 MiB, with its display disabled and a 450 W power limit. The host used a Ryzen 9 7950X, 128 GiB-class system memory with no swap, Ubuntu 22.04.3, NVIDIA driver 570.195.03 and CUDA 11.8. The card was exposed at PCIe 4.0 x8 during capture, so absolute numbers should not be treated as a generic RTX 4090 specification.

The runtime was llama.cpp b10453, exact commit 3cb7ffb1a1f612d5e4a46244ae5a3c77ad934a70, compiled for CUDA architecture 89. Common settings were full GPU offload, flash attention on, F16 K/V cache, batch 2,048, micro-batch 512, 16 threads and one parallel slot. Text lanes did not load a vision projector. Core quality tests kept speculative decoding off.

Two fairness lanes were used:

  • Equal Q4: the three exact Q4_K_M files, the same runtime, F16 K/V and common generation controls.
  • Best usable inside 24GB: configuration changes were allowed only to make a useful context profile fit. The material change was Gemma’s Q8_0 K/V cache at 64K.

Controlled non-thinking tasks used temperature zero and a fixed seed. Controlled reasoning used temperature 1.0, top-p 0.95, medium reasoning effort and three declared seeds. Qwen and Gemma do not expose identical thinking semantics, so we also ran three short native-default prompts per model as non-ranking user-experience samples. Those native samples were not merged into the scorecard.

The deterministic suite contained 50 instruction checks; 40 reasoning cases over three seeds; 30 single-tool and 10 multi-step cases over three seeds; 12 coding tasks over three seeds; 24 document questions across 8K and 32K packs; 12 writing edits; 18 mechanical multilingual checks in Chinese, Spanish and French; and 20 synthetic vision images. Coding agents edited small Python workspaces and ran tests. A claimed success without passing tests remained a failure.

For performance, llama-bench ran five repetitions at 512, 8,192 and 32,768 prompt tokens plus 256-token decode. Context runs required three successful turns; an allocation failure counted as OOM rather than disappearing from the table. Warm streaming tests excluded one warm-up. Binary proportions use Wilson 95% intervals. Paired model comparisons use two-sided exact McNemar tests; p-values describe evidence against symmetric paired outcomes, not practical importance.

One scoring fixture and two exact-match patterns were corrected after inspection: a weekday key was wrong, while valid 63 km and y = 11 answers were too narrowly matched. Retained outputs were rescored; inference was not rerun. Two malformed Qwen3.6 tool calls initially stopped their lanes, so the harness was corrected to record errors and continue. The abort-point attempts stayed as failures, and only genuinely unstarted cases ran afterward.

The evidence ZIP contains prompts, raw responses, tool transcripts, logs, GPU telemetry, commands, scoring scripts and source snapshots. It excludes weights, credentials and private user data. This is one GPU, one llama.cpp commit, one GGUF publisher and a synthetic suite—not a claim about every runtime or workload. The methodology follows Kingy.ai’s guide to how to evaluate Qwen3.8-27B for real work.

Three evidence labels govern the interpretation. Kingy.ai measured means a value was regenerated from a retained run on this host. Derived from pinned configuration means an architectural fact came from the frozen config rather than observed behavior. Vendor-reported means a model-card number that was neither reproduced nor merged with our scores. Confidence intervals are shown because small denominators can make adjacent percentages look more decisive than they are. No composite score was created: weighting one coding task, one vision image and one instruction check equally would manufacture an answer rather than reveal a preference. The fixtures are public and deterministic so readers can inspect their difficulty, but they cannot represent every repository, document, language, image or attack surface. Production selection should repeat the highest-risk tasks with the intended system prompt, tool schemas, sampling policy and failure budget. The workload recommendations below are therefore bounded decisions from this evidence, not rank claims about unrelated uses.

Memory and context: what actually fits in 24GB?

A 16–17 GiB weight file does not leave seven or eight GiB for arbitrary context. The runtime needs working buffers; K/V cache grows with allocated context; projectors and speculative components add their own allocations. “Loads,” “answers one short prompt,” and “is comfortable through a long multi-turn session” are three different states.

Model / equal-Q4 profile 8K peak 32K peak 64K result 64K headroom Interpretation
Qwen3.8-27B, F16 K/V 16,626 MiB 18,186 MiB 20,266 MiB, pass 4,298 MiB (4.20 GiB) Three turns passed at every tested context
Qwen3.6-27B, F16 K/V 16,626 MiB 18,186 MiB 20,266 MiB, pass 4,298 MiB (4.20 GiB) Same measured memory curve as Qwen3.8
Gemma 4 31B-it, F16 K/V 19,962 MiB 21,906 MiB CUDA OOM Not applicable 32K passed with 2,658 MiB free
Gemma 4 31B-it, Q8_0 K/V Not the equal-Q4 cache lane Not required 21,956 MiB, pass 2,608 MiB (2.55 GiB) Best-fit 64K profile; three turns passed

At 32K, Qwen3.8 and Qwen3.6 each peaked at 17.76 GiB and left 6.23 GiB. Gemma peaked at 21.39 GiB and left 2.60 GiB. That is enough to run the tested text session, but a desktop compositor, projector, MTP, parallel slot or larger output reserve can consume the margin quickly.

At 64K, both Qwens passed with F16 K/V and 4.20 GiB free. Gemma’s F16 cache failed during context allocation. Switching only Gemma’s K/V to Q8_0 reduced the cache bill enough for a three-turn pass at 63,482 input tokens, peaking at 21.44 GiB. Its first prefill took 36.57 seconds at 1,756.9 prompt tokens/s; cached follow-up turns were about one second in this synthetic check.

These are highest-tested-stable profiles, not maximum possible context. We did not test 128K or 262K, and the remaining headroom is not a guarantee for every prompt. Context allocation also differs across runtimes. For the wider capacity picture, see the full Qwen3.8-27B memory ladder and why 262K context changes the memory bill.

Measured peak VRAM and headroom at 32K and 64K

Open the full-size memory and context chart.

Speed: prompt processing, decode and end-to-end work

Qwen3.8 and Qwen3.6 were effectively tied in the synthetic throughput lane. Gemma matched them at the shortest prompt, then fell back as context grew and decoded about 8.3% more slowly. Medians below come from five retained repetitions.

Same-GPU measure Qwen3.8-27B Qwen3.6-27B Gemma 4 31B-it
512-token prompt 3,001.35 tok/s 2,980.06 tok/s 2,974.77 tok/s
8,192-token prompt 2,843.75 tok/s 2,840.08 tok/s 2,682.45 tok/s
32,768-token prompt 2,554.36 tok/s 2,554.77 tok/s 2,235.46 tok/s
256-token decode 49.09 tok/s 49.04 tok/s 45.00 tok/s
32K corpus, first-turn characters/s 14,657 14,667 13,650
Warm TTFT, 512-token prompt 0.107 s 0.108 s 0.421 s
Server ready time in TTFT run 24.05 s 22.11 s 25.05 s

Tokenizers were not assumed interchangeable. The character-rate row uses the same long synthetic corpus and helps check that a token/s result is not merely a tokenization artifact. It tells the same story: the Qwens tie; Gemma trails modestly.

The warm streaming probe used a 64-token cap, not a requirement to emit exactly 64 tokens. Its median end-to-end times—0.235, 0.462 and 1.090 seconds—are therefore descriptive only because the models produced different output lengths. TTFT is the cleaner latency comparison. Peak board power during the synthetic benchmark was about 451 W for both Qwens and 454 W for Gemma, close enough that this run does not support an efficiency ranking.

What MTP changed

MTP stayed off for every core quality score. On 15 practical paired prompts, Qwen3.8’s MTP mode produced a 1.64× median end-to-end speedup and 1.66× decode speedup while adding 1,032 MiB peak VRAM. Gemma’s separate 0.479 GiB MTP drafter produced 1.52× end-to-end and 1.67× decode speedups, adding 574 MiB.

The tradeoff is important: Qwen3.8’s baseline and MTP outputs were identical in 0/15 pairs; Gemma’s were identical in 1/15. Separate restart checks showed stable baseline outputs and stable MTP outputs, so the difference was not simple restart noise. MTP is an opt-in speed/behavior tradeoff under this runtime, not a transparent switch. Qwen3.6’s base configuration declares an MTP layer, but this exact Q4 artifact did not contain usable MTP layers and llama.cpp refused the MTP context. That is an artifact/runtime support result, not proof about every Qwen3.6 format.

Same-GPU prompt processing and decode throughput

Open the full-size throughput chart.

Coding, tools and agent reliability

The strongest result in the entire comparison is not a vendor coding score. It is Qwen3.8 completing every one of our 12 code tasks on the first declared seed, then passing 35/36 seeded runs overall. Qwen3.6 and Gemma each passed 16/36 runs, but their case coverage and failure modes differed.

Coding correctness

Each coding case began from a small deterministic Python workspace with a task description and tests. The agent could inspect files, edit code and run tests. Pass@1 uses the first declared seed for each of 12 tasks; pass@3 asks whether any of the three declared seeds solved the task. We retained failed tests, malformed calls and premature-success outcomes.

Qwen3.8 scored 12/12 pass@1 and 12/12 pass@3. Its only failed seeded run was an ordinary test failure. Qwen3.6 scored 8/12 on both measures: retrying with two additional seeds did not expand its solved-task set. Of its 20 failed runs, 19 failed tests and one was the retained malformed-call abort. Gemma moved from 6/12 pass@1 to 7/12 pass@3. Its 20 failed runs included 15 ordinary test failures and five runs with invalid calls; ten invalid-call events occurred in total.

This is strong evidence for Qwen3.8 on these tasks, not a replacement for SWE-bench or a large private repository evaluation. The Wilson 95% interval around 12/12 is still 75.8–100.0%, which is a useful reminder that a perfect small sample has uncertainty.

Tool calling

The declared-tool lane told a more nuanced story. Qwen3.8 and Gemma both passed 90/90 single-call runs. Qwen3.6 passed 79/90; all 11 misses were no-call outcomes rather than calls with wrong arguments. On the multi-step workflows, Gemma passed 30/30, Qwen3.8 28/30 and Qwen3.6 27/30. All three solved every underlying tool case at least once across three seeds.

Gemma therefore deserves the structured-tool recommendation when the interaction is short, schema-bound and closely resembles the declared harness. It does not follow that Gemma is the best coding agent. Its iterative coding runs needed more turns and produced invalid calls that the simpler tool suite did not expose.

Agent efficiency

Among successful coding runs, Qwen3.8’s median model time was 14.28 seconds with five calls. Qwen3.6 needed 24.17 seconds and five calls. Gemma needed 44.39 seconds and ten calls. Those medians exclude failed runs, so they answer “how expensive was a success when one occurred?” rather than masking failures inside an average.

Objective reliability measure Qwen3.8-27B Qwen3.6-27B Gemma 4 31B-it
Instruction checks 43/50 (86.0%; CI 73.8–93.0) 43/50 (86.0%; CI 73.8–93.0) 44/50 (88.0%; CI 76.2–94.4)
Reasoning runs 117/120 (97.5%; CI 92.9–99.1) 28/120 (23.3%; CI 16.7–31.7) 104/120 (86.7%; CI 79.4–91.6)
Single-tool runs 90/90 (100%; CI 95.9–100) 79/90 (87.8%; CI 79.4–93.0) 90/90 (100%; CI 95.9–100)
Multi-step tool runs 28/30 (93.3%; CI 78.7–98.2) 27/30 (90.0%; CI 74.4–96.5) 30/30 (100%; CI 88.6–100)
Coding pass@1 12/12 (100%; CI 75.8–100) 8/12 (66.7%; CI 39.1–86.2) 6/12 (50.0%; CI 25.4–74.6)
Document QA 23/24 (95.8%; CI 79.8–99.3) 8/24 (33.3%; CI 18.0–53.3) 22/24 (91.7%; CI 74.2–97.7)
Vision 18/20 (90.0%; CI 69.9–97.2) 17/20 (85.0%; CI 64.0–94.8) 19/20 (95.0%; CI 76.4–99.1)

The practical conclusion is workload-specific. Gemma was flawless when a tool contract was explicit and short. Qwen3.8 was far more reliable when the agent had to inspect state, edit, test, recover and finish. For local coding agents, completed-task reliability matters more than a pristine one-call schema score.

Reasoning, instruction following and knowledge work

Qwen3.8 and Gemma were both strong on the 40-case reasoning set; Qwen3.8 was more efficient and more reliable within the common 768-token completion budget. Across three seeds, Qwen3.8 passed 117/120 runs and 39/40 cases at pass@3. Gemma passed 104/120 and 38/40. Qwen3.6 passed 28/120 format-compliant runs and 20/40 cases at pass@3; one additional answer was correct but failed the required wrapper.

The failure taxonomy matters more than the headline percentage. Qwen3.8 produced no visible final answer twice and had one wrong answer. Gemma produced no visible final four times and had 12 wrong answers. Qwen3.6 produced no visible final 81 times, with 92 length finishes. Its median completion used the entire 768-token budget and took 16.16 seconds, versus 190 tokens and 4.24 seconds for Qwen3.8, and 522.5 tokens and 12.10 seconds for Gemma.

Several Qwen3.6 traces contained promising or even correct hidden reasoning but exhausted the cap before producing a scorable final answer. The narrow claim is therefore fixed-budget task efficiency, not that Qwen3.6 cannot reason when given more tokens. In a local agent or batch workflow, however, failure to surface the answer inside the allotted budget is operationally real.

Basic non-thinking instruction compliance was nearly identical: 43/50 for each Qwen and 44/50 for Gemma. That result prevents the reasoning gap from being misread as a general inability to follow simple directions. It is specifically the interaction of thinking behavior, output budget and task completion.

Paired outcomes strengthen the Qwen upgrade finding. Against Qwen3.6, Qwen3.8 alone passed 89 reasoning runs; Qwen3.6 alone passed none. The two-sided exact McNemar p-value was approximately 3.23e-27. Against Gemma, Qwen3.8 alone passed 13 runs, Gemma alone none, with p=0.000244. These tests do not measure the size or business value of a difference, but they show that the observed asymmetry is not explained well by random paired flips in this fixture set.

Objective workload reliability with Wilson intervals

Open the full-size reliability chart.

Writing, editing and multilingual use

The mechanical writing checks and the blinded rubric tell different but compatible stories. Mechanical rules rewarded exact word caps, required phrases, sentence counts, bullet structure and number preservation. Qwen3.6 led that lane at 11/12; Gemma scored 9/12 and Qwen3.8 8/12. The paired writing difference between Qwen3.8 and Qwen3.6 was not statistically persuasive in this small set (p=0.25).

For the subjective review, model identities were hidden and 36 outputs were shuffled deterministically. One evaluator scored instruction adherence, factual consistency, style and edit preservation from 0–2. Gemma received 94/96 (97.9%); Qwen3.8 and Qwen3.6 each received 93/96 (96.9%). Three samples only partially removed marketing language, one added an unsupported generalization and one was awkwardly phrased. The one-point spread is too small, with too little evaluator coverage, to support a prose-quality ranking.

The multilingual lane was mechanical rather than a native-speaker quality review. Qwen3.8 and Qwen3.6 each passed 13/18; Gemma passed 14/18. We can say Gemma had the highest retained count, not that it writes best in Chinese, Spanish or French. A real multilingual deployment should add native evaluators, domain vocabulary and locale-specific safety checks.

So did Gemma’s reputation for polished prose appear? Only weakly: it had a one-point blinded edge, while Qwen3.6 led the exact mechanical constraints. For ordinary editing, the responsible verdict is a practical tie; choose using agent reliability, latency, memory and your own voice samples.

Vision, documents and long context

Gemma won the small vision lane by one image: 19/20 versus 18/20 for Qwen3.8 and 17/20 for Qwen3.6. The fixtures covered synthetic charts, flow diagrams, tables and UI screenshots. That makes the result relevant to local OCR-adjacent and interface-understanding work, but the Wilson intervals overlap widely. It does not settle photographic perception, video understanding or real scanned-document quality.

The projectors were loaded only for vision. Qwen3.8 and Qwen3.6 used their exact 0.864 GiB F16 projector files; Gemma used its 1.117 GiB F16 projector. Text speed, memory and quality runs did not pay that cost. If you never send images, do not load a projector merely because the base model is multimodal.

Document QA changed the recommendation. Qwen3.8 answered 23/24 questions and Gemma 22/24. Qwen3.6 answered 8/24; its 16 failures had no visible final answer after consuming the thinking allowance. The two evidence packs targeted approximately 8K and 32K contexts and required grounded answers or citations. Gemma tokenized the same packs more compactly—5,713 and 22,897 tokens versus 6,468 and 26,066 for the Qwens—so token counts alone would not have been a fair workload measure.

All three passed the 32K three-turn context allocation check. Qwen3.8 and Qwen3.6 had much more memory headroom, while Gemma still produced excellent document answers. For a private document assistant, Qwen3.8 is the safest combined choice; Gemma is a close alternative when vision and structured tools matter more. Qwen3.6 may improve with a larger reasoning budget, but that also increases latency and does not repair its smaller measured headroom advantage over Qwen3.8, because their memory curves were identical.

Audio is excluded from this three-way verdict. The tested 31B Gemma checkpoint accepts text and image in our setup, and the Qwen lane used image projectors; no audio model, preprocessing path or audio fixture was run.

Is Qwen3.8 worth upgrading from Qwen3.6?

Yes for coding agents, reasoning and document work. Maybe not for a stable writing-oriented deployment that already meets its targets. The upgrade is unusually attractive because Qwen3.8 did not demand a speed or VRAM sacrifice: it decoded at 49.09 versus 49.04 tok/s, shared the same measured context curve, and its file was only about 0.27 GiB larger.

Upgrade dimension Paired Qwen3.8-only passes Qwen3.6-only passes Ties Exact McNemar p Decision
Reasoning runs 89 0 31/120 3.23e-27 Upgrade
Coding runs 19 0 17/36 0.0000038 Upgrade
Document QA 15 0 9/24 0.000061 Upgrade
Single tools 11 0 79/90 0.00098 Upgrade for tool reliability
Multi-step tools 3 2 25/30 1.0 No clear paired difference
Vision 1 0 19/20 1.0 Too small to decide
Writing checks 0 3 9/12 0.25 Stay if this is your validated niche
Instruction / multilingual 0 0 All items 1.0 Measured tie

Migration is not literally free. The chat templates differ, Qwen3.8 exposes newer reasoning controls, and an application should rerun its own tool schemas, stop conditions and prompt regression suite. Do not copy a Qwen3.6 template into Qwen3.8 and assume equivalence. MTP also differs: it worked from Qwen3.8’s exact artifact under this runtime, while the Qwen3.6 Q4 did not expose usable MTP layers.

The “wait” case is an established Qwen3.6 service whose main job is constrained copyediting or simple non-thinking chat, with no observed completion failures. The “stay” case is a validated production integration where migration risk outweighs the measured task benefit. For a fresh 24GB installation or an agent workflow, the data supports Qwen3.8.

Paired outcome deltas from Qwen3.6 to Qwen3.8

Open the full-size paired-delta chart.

Which model should you choose?

Recommendations should expose the tolerance being optimized, not just name a winner.

User / workload Recommended model Configuration starting point Why Watch for
RTX 3090/4090 coding agent Qwen3.8-27B Q4_K_M, F16 K/V, 8K–32K Best pass@1, seeded completion and successful-run time Validate your real repo and tools; 3090 behavior was not measured here
Private document assistant Qwen3.8-27B Q4_K_M, F16 K/V, 32K 23/24 document QA with 6.23 GiB headroom Retrieval quality and citation policy still matter
Schema-bound tool router Gemma 4 31B-it Q4_K_M, F16 K/V, 8K Perfect single and multi-step declared-tool lane Iterative coding calls were much less reliable
Vision-heavy assistant Gemma 4 31B-it Q4_K_M plus F16 projector Highest vision count, strong documents Larger projector and less text-context headroom
Writer or role/chat user Test all three Native template and your own style set Blinded writing result was effectively tied Mechanical constraints and voice preference differed
64K text workspace Qwen3.8-27B Q4_K_M, F16 K/V Passed with 4.20 GiB headroom 64K was highest tested, not guaranteed maximum
Existing stable Qwen3.6 stack Qwen3.6-27B or staged upgrade Keep current profile until regression tests pass Avoids migration risk Large measured deficits in agents, reasoning and documents
Speed above dense-model quality Compare a MoE separately Outside this test Qwen3.6-35B-A3B / Gemma 4 26B-A4B are adjacent candidates They were not tested here and are not declared winners

Nominal 24GB capacity does not make every card equivalent. An RTX 3090 has different bandwidth, power and software behavior from this RTX 4090; workstation cards may reserve memory differently. Treat the configuration as a starting point, then measure your actual prompt lengths, output reserve, concurrency and desktop overhead.

What vendor benchmarks do—and do not—show

Vendor model cards are valuable for hypothesis generation. They are not interchangeable with this local Q4 run. Harnesses, prompts, checkpoint formats, token budgets, benchmark versions and reasoning settings differ—even when a row carries the same benchmark name.

Vendor-reported source context Benchmark Qwen3.8 Qwen3.6 Gemma 4 What it suggests, not proves
Qwen3.8 model card Terminal Bench 2.1 (Terminus) 73.0 63.4 Qwen claims stronger terminal-agent work
Qwen3.8 model card SWE-bench Pro 61.7 53.5 Qwen claims an upgrade in software tasks
Qwen3.8 model card IFBench 79.5 69.1 Qwen claims better instruction following
Qwen3.8 model card OSWorld-Verified 84.3 63.9 Qwen claims a larger computer-use gain
Qwen3.8 model card OmniDocBench 1.5 91.1 89.4 Qwen claims a smaller document gain
Qwen3.6 model card SWE-bench Pro 53.5 35.7 Its table favors Qwen3.6 over Gemma
Qwen3.6 model card MMMU-Pro 75.8 76.9 Its table gives Gemma a slight vision/reasoning edge
Gemma 4 model card AIME 2026, no tools 89.2 Gemma reports strong mathematical reasoning

The source records are the pinned Qwen3.8 model card, Qwen3.6 model card and Gemma 4 model card. The table deliberately does not merge those values with Kingy.ai measurements or construct a combined score.

Our local findings broadly agree with the Qwen3.8 upgrade direction for coding, agents and documents, while showing that Gemma can be excellent at structured tools and vision. Agreement does not validate the vendor harnesses; disagreement would not automatically invalidate them. They answer different questions.

Final verdict

Qwen3.8-27B is the best all-round choice in this exact 24GB comparison. It combined Qwen3.6’s speed and memory behavior with a decisive improvement in coding completion, fixed-budget reasoning, document QA and single-tool reliability. For a new RTX 3090/4090-class local agent or document system, start there.

Gemma 4 31B-it is not an also-ran. Choose it when declared tools and vision dominate, and accept its higher VRAM use, slower text throughput and weaker iterative coding result. The blinded writing test was essentially tied, so prose preference should be settled with your own samples.

Qwen3.6-27B remains defensible as a compatibility choice, not the strongest fresh deployment. Its Q4 file is slightly smaller, but it offered no practical speed or context advantage over Qwen3.8 on this host.

The most important 24GB constraint is the whole allocation, not the weight file. Qwen3.8 and Qwen3.6 kept 4.20 GiB free at the highest stable 64K profile tested; Gemma required Q8_0 K/V and kept 2.55 GiB. Add projectors or MTP only when the workload benefits, and remeasure headroom after every change.

FAQ

Is Qwen3.8-27B better than Qwen3.6-27B?

For the tested coding-agent, reasoning, document and single-tool workloads, yes. Qwen3.8 matched Qwen3.6’s 49 tok/s decode speed and memory curve but produced 19 additional coding-run passes, 89 additional reasoning-run passes and 15 additional document passes in paired outcomes. Basic instruction and multilingual mechanical checks tied, while Qwen3.6 led three small writing checks. Existing deployments should still regression-test the new template.

Is Qwen3.8-27B better than Gemma 4 31B?

It depends on the job. Qwen3.8 was clearly stronger in our coding harness, more accurate and efficient in controlled reasoning, slightly better in document QA, faster in text generation and lighter in VRAM. Gemma was perfect in the structured tool lane, scored 19/20 in vision and edged the single-evaluator writing rubric by one point. Qwen3.8 is the safer general default; Gemma is the specialist alternative.

What is the best local LLM for a 24GB GPU?

Among these three dense multimodal Q4 models, Qwen3.8-27B is the best overall starting point. That does not cover every local model: a speed-first MoE, smaller model, CPU-offloaded larger model or domain fine-tune may fit another priority. “Best” should be defined by completed tasks, latency, context and headroom on your own hardware—not parameter count or a single benchmark row.

Can Qwen3.8-27B run fully in 24GB VRAM?

Yes, the tested 15.932 GiB Q4_K_M artifact ran with full GPU offload. It passed three-turn allocations at 8K, 16K, 32K and 64K using F16 K/V, peaking at 20,266 MiB at 64K and leaving 4,298 MiB. That does not guarantee 128K/262K, concurrency, a loaded projector or MTP will fit simultaneously. Reserve output and desktop overhead before calling a profile comfortable.

Can Gemma 4 31B run on an RTX 3090 or RTX 4090?

The exact Gemma 4 31B-it Q4_K_M ran fully on this RTX 4090. It passed 32K with F16 K/V at 21,906 MiB. Its 64K F16 cache was OOM, while Q8_0 K/V passed at 21,956 MiB. An RTX 3090 also has 24GB, but this test did not measure it; expect different throughput and verify driver, display reservation and real headroom locally.

Which model is best for local coding and tool use?

Choose Qwen3.8 for an agent that must inspect, edit, test and recover: it scored 12/12 coding pass@1 and 35/36 seeded runs, with zero invalid calls. Choose Gemma for short, tightly declared tool contracts: it passed all 120 single and multi-step runs. Qwen3.6 was intermediate in coding case coverage but had lower single-tool reliability and a retained malformed-call failure.

Which model is best for writing and general chat?

This test found no meaningful prose winner. Gemma scored 94/96 in a model-blinded, single-evaluator rubric; both Qwens scored 93/96. Qwen3.6 led the mechanical constraint checks at 11/12. The samples were short edits, not long fiction, roleplay or a diverse reader panel. Test native templates with your tone, refusal requirements and preferred response length before choosing for chat.

How much context can each model use on 24GB?

The highest stable context tested was 64K. Qwen3.8 and Qwen3.6 passed three turns with F16 K/V and 4.20 GiB headroom. Gemma failed 64K with F16 K/V but passed with Q8_0 K/V and 2.55 GiB headroom. All three configurations advertise up to 262K, but advertised position capacity is not the same as a practical 24GB allocation.

Should I use Q4_K_M, Q5 or another quant?

Q4_K_M is the defensible baseline here because all three exact files fit fully and can be compared consistently. A Q5 may improve quality but reduces context and optional-component headroom; smaller quants trade quality for capacity. Do not infer a Q5 winner from these Q4 results. Select the context and output reserve first, then test the highest quant that leaves a safety margin under your real runtime.

Does MTP improve Qwen3.8 speed?

Yes, under the pinned llama.cpp build. Qwen3.8’s practical MTP pairs showed a 1.64× median end-to-end speedup and 1.66× decode speedup, at an extra 1,032 MiB peak VRAM. However, all 15 MTP outputs differed from their baselines. Treat it as a behavior-changing optimization that needs task-level validation, not a free speed toggle. Core quality results in this article kept MTP off.

Do I need the vision projector?

Only for image input. Text-only inference worked without it and all text speed, quality and memory lanes omitted it. The exact F16 projector files were 0.864 GiB for each Qwen and 1.117 GiB for Gemma. Load the matching projector revision when testing charts, screenshots or photos, then remeasure VRAM. A mismatched projector can invalidate vision results even if the text model loads.

Are Qwen3.8 and Gemma 4 open source?

Their pinned model cards list Apache 2.0 and provide downloadable weights; Qwen3.6 does as well. We still use “open-weight” in this comparison because “open source AI” can imply broader disclosure of training data, recipes and reproducibility. Check the Qwen3.8 license record and Gemma 4 license terms for your use case rather than relying on shorthand.

Sources, test manifest and change log

Evidence cutoff: 2026-08-17 06:31:33 UTC. Physical test date: August 16, 2026, Pacific time. Canonical model records, configs, generation configs and chat templates were refreshed at the cutoff and retained with response headers and SHA-256 hashes. GGUF hashes matched the predeclared artifact manifest before testing. The llama.cpp release record and exact commit are included.

Download the 24GB comparison evidence bundle (ZIP, approximately 4.6 MB; SHA-256 a6a793932b75ce58836b2053042aec18b08de211913fee130d499b3d196ba0cd). It contains 2,110 files: deterministic fixtures, retained outputs, coding workspaces, server summaries, GPU telemetry, OOM logs, source snapshots, the blinded writing packet and key, scoring scripts, regenerated statistics, a manifest and per-file checksums. Model weights, binaries, credentials and private data are excluded.

Primary records: Qwen3.8-27B, Qwen3.6-27B, Gemma 4 31B-it and llama.cpp b10453. Community discussions were used only to understand search intent, never as support for measured memory, speed or quality claims.

Change log:

  • 2026-08-16: Initial same-hardware Q4 comparison, source audit, raw evidence bundle and article publication candidate.