Qwen 3.8 27B has the same architecture as 3.6, line for line: 100% of the gain came from training
Alibaba shipped a new model and the architecture diff came back empty
Qwen dropped the weights for Qwen 3.8 27B on August 14, 2026, at 15:00 UTC — midnight on the 15th in Japan. Apache 2.0, dense, natively vision-language. The usual ritual: benchmark table, big jump, everyone sprinting to Ollama.
Then someone did the most boring and most useful thing possible: opened the 3.6 config and the 3.8 config side by side.
Nothing changed. Not a layer, not a head, not a dim. The August model's graph is the same one from April, and the benchmark jump is huge. This post is about what that proves — and what it does not prove.
TL;DR
- What it is: a dense model with 27.78 billion parameters, natively vision-language, 262K context, Apache 2.0.
- The finding: the architecture diff against 3.6 27B has exactly one differing field, and it's metadata (
transformers_version). 59 of 59 graph nodes map one-to-one. - What it means: the gain did not come from architecture. It came from training, post-training and optimization — none of the three disclosed.
- Where it hurts: a KV cache of ~64KB/token,
xhighas the default reasoning effort, and the official Jinja template breaking OpenAI tool calls. - Cost/access: open weights on Hugging Face; via API, $0.45 input / $3.20 output per 1M on OpenRouter.
- Don't mix them up: Qwen 3.8 27B is the open-weights dense model that runs on your machine. Qwen 3.8 Max is a different model, the 2.4T-parameter MoE with ~95B active, served over API. Two launches, two different decisions.
The diff that proved it: same architecture as 3.6 (and 3.5)
The comparison is public on hfviewer, and its wording is dry: "59 of 59 source-faithful nodes map one-to-one; no nodes, edges, repeat counts, or graph-visible dimensions changed."
In practice, both configs say the same thing:
// Qwen3.6-27B (abril/2026) e Qwen3.8-27B (agosto/2026)
"architectures": ["Qwen3_5ForConditionalGeneration"],
"hidden_size": 5120,
"num_hidden_layers": 64,
"intermediate_size": 17408,
"vocab_size": 248320,
"max_position_embeddings": 262144
// padrão híbrido: 16 × [3 × (Gated DeltaNet → FFN) + 1 × (Gated Attention → FFN)]
// 24 heads Q / 4 KV @ dim 256 · 48 heads V / 16 QK @ dim 128
//
// único campo com valor diferente entre os dois: "transformers_version"
Look at the class name: Qwen3_5ForConditionalGeneration. Not a typo. That's the 3.5 generation block still declared in the 3.8 file. Three releases, one design.
Now the part the hype skipped, written on the same page as the finding: "This is an architecture result, not a claim that the checkpoints are identical. Training data, optimization, learned weights, post-training and behavior can change substantially while the computation graph stays fixed."
In other words: the weights are different. What's identical is the blueprint of the house, not the house. Anyone going around saying "it's byte-for-byte the same as 3.6" will get corrected by the first person who opens the safetensors.
And, to be honest all the way through: the evidence only supports "it wasn't architecture". The community filled in "therefore, it was data" on its own. It could have been data, it could have been post-training, it could have been the optimization recipe — Qwen disclosed none of the three, nor token counts, nor knowledge cutoff, nor a safety evaluation. We traded an architecture black box for a training black box.
None of that, by the way, is what's going to burn you. What burns you is the Tuesday your agent dies with a Jinja template error and you lose the whole day hunting for the fix in some buried Hugging Face discussion — the kind of thing that gets solved in minutes when you have people to compare notes with, and in days when you're on your own. That's basically what the Clã Beer and Code is: people building with AI live, every week, sharing what has already broken.
Real benchmarks, and what the vendor table leaves out
The numbers for 3.8 27B against 3.6 27B, all reported by Qwen itself:
| Benchmark | 3.6 27B | 3.8 27B |
|---|---|---|
| Terminal-Bench 2.1 | 63.4 | 73.0 |
| SWE-bench Pro | 53.5 | 61.7 |
| QwenSWEBench | 49.3 | 79.0 |
| DeepSWE 1.1 | 13.3 | 42.2 |
| OSWorld-Verified | 63.9 | 84.3 |
| Agents' Last Exam (pass@1) | 10.6 | 20.4 |
| GPQA Diamond | 87.8 | 89.2 |
DeepSWE 1.1 is the brutal number: it tripled on the same graph. And OSWorld-Verified, +20.4 points, is computer use — the model is natively vision-language, not text-only. Same network design, twenty more points driving a screen.
Another data point published without fanfare: the dense 27B beats Qwen3.7-Plus, the higher tier in Qwen's own lineup, on SWE-bench Pro (61.7 vs 57.6), QwenSWEBench (79.0 vs 59.2), CoWorkBench (70.7 vs 61.0) and JobBench (33.4 vs 21.8). A model that runs on your machine beating the paid one from the same shop.
Against Muse Glimmer 30B, the press stamped it with "beats on many benchmarks" — Terminal-Bench 73.0 vs 51.7, GPQA Diamond 89.2 vs 83.5. Except the table leaves Glimmer out on NL2Repo-Bench, DeepSWE 1.1, JobBench and LiveCodeBench v6. An incomplete comparison by vendor omission.
And where it loses: 5.2 points behind Opus 4.6 Max on Terminal-Bench (73.0 vs 78.2), 9.2 on HLE, and NL2Repo-Bench 42.3 vs 47.6.
Methodology warning, and it's worth half the post: zero of these numbers have been independently reproduced. kingy.ai is literal about it — "Every launch score comes from Qwen", "several benchmarks are in-house, corrected or modified". Even "FP8 ties BF16" is, in their words, "a vendor claim until evaluated".
The thinking token problem: 10x more for the same result
This is where the bill that doesn't show up in the table lives.
The card recommends xhigh as the default reasoning effort, and preserve_thinking ships on by default across all workloads. Put the two together: a model that thinks a lot, thinks again, and carries its reasoning forward between turns.
In the Hacker News thread — 1,269 points and 741 comments when I checked, a number that keeps moving — CMay says 3.8 got a private benchmark of his right, but "it took 5x as many tokens to do it and 12m30s with MTP enabled". Casteil is harsher: Gemma 4 26B gets to the same answer with "only 1/10th as many 'thinking' tokens", in a fraction of the time. And dofm describes the trace as "almost caveman" — the model drops function words and writes in note form ("Need be helpful concise"), raising the hypothesis that "this rather unique thinking trace pattern is actually hobbling the MTP predictions".
Translated to product terms: if the model burns 5x more reasoning tokens for the same result, the price per 1M output tokens is no longer the metric that matters. Turn the effort down.
{
"model": "qwen/qwen3.8-27b",
"reasoning": { "effort": "low" },
"messages": [{ "role": "user", "content": "..." }]
}
Do not just follow the news — master it. Hands-on AI Engineering, live, every week, in the largest community in Brazil.
Join the ClãRunning Qwen 3.8 27B locally: real VRAM, KV cache and the Jinja template fix
The press boiled the requirement down to "24GB, RTX 4090 class". The real numbers are different:
| Format | Weights | With native KV |
|---|---|---|
| BF16 | 51.76 GiB | ≥ 67.76 GiB (80GB+ system) |
| FP8 | 28.76 GiB | ≥ 44.76 GiB (48GB minimum) |
| Q4_K_M GGUF | 15.93 GiB | +2 to 4 GiB at 32–64K |
The line from the source: "A 24GB GPU is a plausible target for a four-bit quant at moderate context, not for BF16, FP8 or the full 262K window." On a 32GB RTX 5090 you run Q4 comfortably and FP8 never.
But the real cost isn't the weights — it's the KV cache. The math going around HN puts it at ~64KB per token on Qwen versus ~13KB on Glimmer: 32K of context eats 2.5GB of VRAM, while Gemma 4 and Glimmer deliver 256k–768k in the same memory, with Glimmer hitting "4x the aggregate tokens/s". And quantizing the cache to compensate doesn't save you: the report is that "doing any quantizing definitely hurt results a lot" — the model demands F16 on K and V.
The second stumble is the template. The official 3.8 one has four known bugs, two of them blockers for anyone building an agent: it "throws a fatal runtime exception if you pass enable_thinking=false" — you can't turn reasoning off — and it breaks with TypeError: Can only get item pairs from a mapping when it receives tool arguments in OpenAI format. In the Unsloth repo, troed reported "System message must be at the beginning." running GGUF on llama.cpp; danielhanchen fixed it in ~23h.
The community fix is a single file, froggeric/Qwen-Fixed-Chat-Templates v22 (August 13, 2026), covering 3.5, 3.6 and 3.8:
llama-server -m Qwen3.8-27B-Q4_K_M.gguf \
--jinja \
--chat-template-file chat_template.jinja \
--reasoning-format deepseek \
--ctx-size 32768 \
--cache-type-k f16 --cache-type-v f16
If you've never stood up a local model end to end, the path is spelled out in our guide on how to run an LLM locally.
Qwen 3.8 27B dense vs Qwen 3.8 Max (2.4T MoE): which one is for you
A common mix-up this week, and worth untangling: these are two different models.
3.8 27B is the dense model in this post: 27.78B parameters, open weights, runs on your machine. Qwen3.8-Max is the top-of-the-line MoE (Qwen3.8-2.4T-A95B) — 2.4 trillion total parameters, ~95B active, announced on August 3, 2026, and with its own post here: Qwen 3.8 Max, plus the follow-up on the open weights.
The practical call: dense 27B if the requirement is data that never leaves your infra, predictable cost or latency without a network hop — via API, $0.45/$3.20 per 1M versus $2/$6 for Max, ~4.4x cheaper on input. Max if the task is frontier-grade and you're fine depending on a third-party API.
It's not "smaller vs bigger". It's "where the data lives vs how much reasoning you need".
Limitations and things to watch
Where this release is still in the dark, no sugarcoating:
- No independent reproduction. Every score comes from Qwen. Don't treat the table as engineering fact until someone outside runs it.
- Training undisclosed. Data, token counts, knowledge cutoff, safety evaluation: nothing published. The "it was data" thesis is inference, not documentation.
- Community tokens/s aren't comparable. Reports range from 40 to 200 tokens/s depending on quantization, speculative decoding and hardware.
- Loading is not operating. NxCode: "A laptop that loads the model may still struggle with a 256K agent session." With
preserve_thinkingon by default, you also carry a wrong assumption from one turn into the next. - Unconfirmed rumors: a reported sighting of a "Qwen 3.8 35B-A3B" (the published 35B-A3B is from the 3.6 generation) and an abliterated version of 3.8 (the ones I found are all 3.6). Numbers from Reddit threads floating around weren't verified here either — the only thing actually verified is the Hacker News thread linked above.
- Security: open weights with no published safety evaluation don't become a customer-facing endpoint without your own guardrails and logging.
FAQ
If the architecture is the same, can I use 3.8 as a drop-in for 3.6?
At inference, almost — the graph is the same and the serving stack recognizes it. In practice no, because of the template: the official 3.8 one breaks with enable_thinking=false and with OpenAI tool calls. Swap the template before you point your agent at it.
Does it run on my 32GB RTX 5090? Q4_K_M yes, with room to spare (15.93 GiB of weights + KV). FP8 doesn't fit with usable context, and BF16 isn't even close. Large context costs more VRAM than the model itself.
Is it better than Gemma 4 26B for a local agent? Depends on what hurts you more. Qwen scores higher on agent benchmarks; Gemma 4 solves similar tasks with a fraction of the reasoning tokens and a lot more context in the same VRAM. Measure it on your case, with your prompt.
Why does enable_thinking=false throw an exception?
A bug in the official Jinja template, not in the model. Fixed in froggeric's package and in Unsloth's UD-* variants.
Conclusion
The most interesting thing about this release isn't the model. It's the public evidence that, in four months, the same network — same graph, same dims, same class inherited from 3.5 — tripled on a software engineering benchmark. Architecture stopped being the bottleneck. Training and post-training are the battlefield now.
Just don't confuse "it wasn't architecture" with "it was data". The first statement is proven by a diff. The second is a polite guess about a black box.
And on your machine, none of that weighs as much as the KV cache, the broken template and reasoning set to the wrong default. A good model you can't operate doesn't become a product. Lower the reasoning_effort, swap the template, measure tokens per task — and only then look at the table.
{AI Engineer} — apaixonado por Laravel, arquitetura de software e construir produtos com impacto. Compartilho aqui tutoriais, descobertas e reflexões sobre o dia a dia de engenharia.
There is no shortage of content. What is missing is someone to untangle it: what matters now is how to implement it the right way. In the Clã you get that live, every week, with people who have already filtered out the noise.
Join the Clã