~/beer-and-code
▪ next event Workshop: Jev na Prática para Devs · 14 Oct · 19h — duração de 2h a 3h, ao vivo via Google Meet save your seat ›
~ / news / qwen-3-8-max-open-weights $
News

Qwen 3.8 Max open weights on the 10th: what you can (and can't) run out of 2.4 trillion

LS Lucas Souza · · 12 min read
Qwen 3.8 Max open weights on the 10th: what you can (and can't) run out of 2.4 trillion

On Monday, August 3, Alibaba announced Qwen3.8-Max and promised something it had never done before: publishing the weights of a Max-class model. The open-weights Qwen 3.8 Max lands "next week" — meaning starting Monday, August 10, on Hugging Face and ModelScope.

That's 2.4 trillion parameters. Mixture-of-Experts, ~95 billion activated per token, 1 million tokens of context, native text, image and video input. By the company's own numbers, it's the most capable model Qwen has ever trained.

And then comes the question every dev asks in the thread and almost nobody answers with a number: can you run it? This post pulls apart the three things being treated as one — open weights, permissive license and viable hardware. Spoiler: only one of them is confirmed, and it's not the one you want.

TL;DR

  • What it is: Qwen3.8-Max, a MoE with 2.4T total parameters and ~95B activated, 1M-token context (991K max input, 131K output), multimodal on input. Announced on August 3, 2026.
  • Open weights: promised for the week of August 10 on Hugging Face and ModelScope. It's the first time Qwen opens a checkpoint from the Max line.
  • Qwen3.8-27B ships alongside it: the little sibling, also open-weights. That's the one most people will actually be able to run.
  • License: not announced. Qwen 3.5 and 3.6 shipped under Apache 2.0, but nothing guarantees the Max follows the same path.
  • API today: $2 per million input tokens, $6 on output, $0.25 with implicit caching. Available through Alibaba Cloud Model Studio.

Open-weights Qwen 3.8 Max: why this changes the game

Until now Qwen operated with a clear split: the open line (7B, 27B, 235B) pulled the open source ecosystem along, and the Max line stayed closed, behind the API, as a commercial product. Every big Chinese lab works this way. Moonshot was the first to break the pattern by publishing Kimi K3 with 2.8 trillion parameters on July 26. Two weeks later, Alibaba answers with its own.

This isn't a release-notes detail. It's a shift in competitive strategy: when the top of the line becomes a distributable commodity, the lab's differentiator stops being the model and becomes the infrastructure for serving that model. The ones who profit are inference providers and the companies that already know how to operate a cluster. The one for whom nothing changes is the dev who thought "open" meant "fits on my machine".

And that's exactly where almost everyone misreads it. Open weights answer the permission question. They don't answer the capacity question — and the distance between the two is measured in terabytes of VRAM and in a license clause nobody has read yet. Reading a model release by separating those layers is an engineering skill, not a timeline-following one: it's the kind of math we do together every week, live, in the Clã Beer and Code, with numbers on the table instead of benchmark screenshots. It's paid, it's a subscription, and it's the environment this post describes.

What Alibaba says the model does (and the asterisk)

When the model was unveiled in July, we had already dug into what was real in the first batch of Qwen 3.8 Max benchmarks. Now the official table is out, and the published numbers are strong:

Benchmark Qwen3.8-Max Comparison
Terminal-Bench 2.1 86.6 GPT-5.6 Sol (max): 88.8 · Claude Opus 4.8 / Fable 5: 84.6
PaperBench 93.0 highest score reported to date
GPQA Diamond 92.6 —
OSWorld-Verified 86.1 Claude Fable 5: 85.0 · GPT-5.6 Sol Max: 83.2 · Gemini 3.1 Pro: 76.2
IFBench 82.8 —

Now the asterisk, which matters more than the table: every one of these scores comes from Alibaba's internal runs. There's no independent verification published so far. The only third-party read that has surfaced puts the model tied with Claude Opus 4.7 on the Vals Index, at a cost per test roughly 2.3x lower — which is great, and is a different story from "beats GPT-5.6".

Treat the table as what it is: the vendor describing its own product. The honest conclusion left standing is the cost one. At $2 / $6 per million tokens, Qwen3.8-Max delivers frontier performance at a third-tier price. That's enough to put it in your decision matrix without depending on any benchmark being true.

Running Qwen 3.8 Max locally: the hardware math on 2.4 trillion

Here's the point the headline swallows. A 2.4T MoE activates ~95B per token — meaning the compute cost per token is that of a 95-billion dense model, which is perfectly reasonable. The problem isn't FLOPs. It's memory: all 2.4 trillion have to be loaded, because the router can call any expert at any layer.

How much that weighs depends on the format Alibaba publishes — and that hasn't been said yet:

Weight format Approximate size
BF16 (2 bytes/param) ~4.8 TB
FP8 (1 byte/param) ~2.4 TB
FP4 / MXFP4 (0.5 byte/param) ~1.2 TB

Now cross that with hardware that actually exists, keeping in mind that KV cache, activations and runtime overhead still go on top:

Configuration Aggregate VRAM Does it fit?
1× RTX 5090 (32 GB) 0.032 TB No, and the joke isn't even funny
8× H100 (80 GB) 0.64 TB No, in no format
8× H200 (141 GB) 1.128 TB No — you're short before you start, even in FP4
8× B200 (180 GB) 1.44 TB FP4 only, with no headroom for context
8× B300 / MI355X (288 GB) 2.304 TB Fits in FP8, tight
GB300 NVL72 ~20.7 TB With room to spare

Look at the H200 row. It's the same trap as K3: 8 × 141 GB looks like a lot until you compare it against the right number. It doesn't fit.

That's why reports from people who have already geared up to serve the model talk about a minimum of 8 B200-class accelerators, with providers recommending supernode configurations of 64+ GPUs to handle real load with long context. This is not "runs on my home setup with aggressive offload". It's a datacenter project, with datacenter interconnect — and in a MoE, bad interconnect hurts a lot more than in a dense model, because token routing crosses the network at every layer.

If you want the full version of this math, with the break-even of renting versus API, it's in the Kimi K3 post. The arithmetic doesn't change its name just because the model changed its logo.

▪ Clã Beer and Code

Do not just follow the news — master it. Hands-on AI Engineering, live, every week, in the largest community in Brazil.

Join the Clã

Qwen3.8-27B is the model that matters to you

Buried in the same announcement is the checkpoint almost all the coverage treated as a footnote: Qwen3.8-27B, also going open-weights on the 10th.

Twenty-seven billion parameters, dense. Here the math gets human:

Precision Weights + KV cache Total GPU that handles it
BF16 ~54 GB ~12 GB ~66 GB 1× H100 80GB
FP8 ~27 GB ~12 GB ~39 GB 1× L40S 48GB
INT4 ~15 GB ~12 GB ~27 GB 1× RTX 5090 32GB

Two honest caveats on those numbers. First: they're a third-party estimate from the parameter count, because Alibaba has not published a spec sheet or benchmarks for the 27B — the model was announced, not documented. Second, and this one is a real production trap: the KV cache doesn't shrink when you quantize the weights. Going from BF16 to INT4 cuts 39 GB of weights and zero bytes of cache. If your plan is long context, the cache becomes the bottleneck and you need to quantize it separately.

A rule of thumb that holds for any model in this range: INT4 is great for classification, extraction and routing — single-step tasks, where losing a point of precision costs little. For multi-step agentic work, stay at FP8 or above. Compounding error across a chain of 20 calls turns a small accuracy degradation into a big drop in completed tasks.

For the step-by-step of getting this up on your machine, we have the guide on how to run a local LLM with Ollama and the hardware discussion.

Qwen 3.8 Max license: the detail that decides everything (and doesn't exist yet)

This is the hole in the middle of the announcement, and it's what I'd look at first on the 10th.

Alibaba has not announced which license the weights will ship under. Not for the Max, not for the 27B. There's no published text to read. Qwen 3.5 and Qwen 3.6 shipped under Apache 2.0, which sets a reasonable expectation — but an expectation isn't a clause, and a Max-class frontier checkpoint is exactly the kind of asset that makes a lab's legal team invent custom terms.

The spectrum of possibilities runs from:

  • Apache 2.0 — free commercial use, no revenue cap, no attribution requirement in the interface. Best case.
  • Modified MIT (what Moonshot did with K3) — commercial use allowed, with triggers above a certain revenue or user count.
  • Custom community license — acceptable-use restrictions, territory clause, attribution requirement, a ban on training other models with the outputs.

The difference between the first and third options isn't philosophical. It decides whether the model goes into your product or not. Readings have circulated in the community that there would be usage restrictions in the US, European Union, United Kingdom and Korea — with no confirmation from Alibaba and no license text to check. Don't make an architecture decision on top of that before the 10th.

The checklist for the day the weights drop, in order:

  1. Open the repository's LICENSE before looking at any benchmark.
  2. Look for a revenue trigger, user cap, attribution requirement and geographic restriction.
  3. Check the published format (BF16, FP8, FP4) — that's what defines the VRAM table up above.
  4. Only then evaluate whether the model goes into your matrix.
# no dia 10, comece por aqui — o LICENSE antes do model card
hf download Qwen/Qwen3.8-27B LICENSE --local-dir ./qwen3.8-27b

# baixar o 27B completo (formato e nome exato do repo confirmam no dia)
hf download Qwen/Qwen3.8-27B --local-dir ./qwen3.8-27b --resume

# servir com vLLM, FP8, contexto conservador
vllm serve ./qwen3.8-27b \
  --quantization fp8 \
  --max-model-len 32768 \
  --gpu-memory-utilization 0.90

The exact repository name is only confirmed once they publish. Treat the commands above as a shape, not as copy and paste.

Limitations and things to watch

"Next week" is not a date. Alibaba said "next week", the press rounded it to August 10. There's no official date with a day and time. We've seen a lab deliver ahead of schedule (Moonshot dropped K3 a day early) and we've seen labs slip by weeks. Don't build a sprint schedule on top of this.

1M context is the advertised ceiling, not the operational default. Every token in the window takes up memory that competes with the model weights on the same node. In production you'll set a much lower per-request limit, and you'll be happier with the result.

Multimodal in, text out. The model accepts image and video, but returns text. If your use case calls for image generation, this isn't the model.

A benchmark is not your domain. Always true, doubly true when the numbers are self-reported. Without your own eval suite on top of your own data, "Qwen3.8-Max is better" is an opinion with a pretty table.

Quick FAQ

Can I run Qwen3.8-Max on my machine with aggressive quantization? No. Even in FP4, the weights come in around 1.2 TB — about 37x the VRAM of an RTX 5090. Offloading to RAM and disk technically "runs", in the sense that you see a token show up every now and then. For real local use, the target is the 27B.

What's the Qwen 3.8 Max license? It hasn't been announced yet, not for the Max and not for the 27B. Qwen 3.5 and 3.6 used Apache 2.0, which is a signal, not a guarantee. The only way to answer this is to read the LICENSE file once the weights are up.

Is the API or self-hosting the better deal? At $2 / $6 per million tokens, the API wins for practically everyone on cost. Self-hosting the Max only pencils out for other reasons: data sovereignty, a regulatory requirement that forbids inference from leaving the perimeter, or fine-tuning on the weights. If the argument is savings, redo the spreadsheet.

What changes for people running Qwen 3.6 locally today? Probably not much in the short term, because the new 27B shipped with no published benchmark. The rational migration is to wait for the first independent numbers, run your own evals against 3.6 and only switch if it wins in your domain. A newer model isn't automatically better for your case.

Conclusion

Qwen3.8-Max going open is big, real news: it's the first time a lab of this scale publishes the top of its own line, and it confirms the move Kimi K3 opened two weeks earlier. The open-weights ceiling just went up again.

What didn't go up is the number of people who can use it. It's still a datacenter artifact, with a license nobody has read, benchmarks nobody has verified and a 27B sibling the coverage treated as a footnote that is, in practice, the only one of the two that will run on your infra.

The read worth carrying into the 10th is this: when the frontier model becomes a public file, the competitive advantage moves entirely to whoever knows how to operate it. The model became a commodity. Engineering didn't.

Lucas Souza
Written by
Lucas Souza

{AI Engineer} — apaixonado por Laravel, arquitetura de software e construir produtos com impacto. Compartilho aqui tutoriais, descobertas e reflexões sobre o dia a dia de engenharia.

▪ Clã Beer and Code

There is no shortage of content. What is missing is someone to untangle it: what matters now is how to implement it the right way. In the Clã you get that live, every week, with people who have already filtered out the noise.

Join the Clã
Meet the Clã Beer and Code
playing