~/beer-and-code
▪ next event Workshop: Jev na Prática para Devs · 14 Oct · 19h — duração de 2h a 3h, ao vivo via Google Meet save your seat ›
~ / news / qwen-3-8-max $
News

Qwen 3.8 Max (2.4T): Benchmark vs Kimi K3, Pricing and Open Weights

LS Lucas Souza · · 9 min read
Qwen 3.8 Max (2.4T): Benchmark vs Kimi K3, Pricing and Open Weights

Alibaba dropped Qwen 3.8 Max on stage at the World AI Conference in Shanghai, and the effect was immediate: half the internet shouting that it's the "open source Fable," the other half calling it a marketing stumble. In the middle sits a detail nobody should skip: on announcement day, there was no benchmark table, no model card and no license.

The model is big. 2.4 trillion parameters, multimodal, a 1-million-token window. Alibaba says it's the second-strongest model in the world, behind only Anthropic's Fable 5. Except that score came from an internal evaluation. Nobody outside has run it.

In this post we separate fact from narrative: what Qwen 3.8 Max actually is, what the community tested in the first 24 hours, and how you can access it today without getting it mixed up with the old qwen3:8b.

TL;DR

  • What it is: Qwen3.8-Max-Preview, Alibaba's 2.4T-parameter multimodal model (sparse MoE), announced on 07/19/2026 at WAIC in Shanghai.
  • Modalities: text, image, video and document. 1M-token context. API compatible with the OpenAI and Anthropic protocols.
  • Benchmarks: none public. The "second only to Fable 5" claim is an internal Alibaba eval, with no table and no methodology.
  • Access/cost: preview via Token Plan, Qoder and QoderWork, at 10% of the standard price. Open weights promised, with no date and no license.

Update: since this post, Alibaba confirmed the ~95B active parameters and announced the open weights for Max, which we broke down in Qwen 3.8 Max open weights: what you can (and can't) run. The checkpoint that runs on your machine is Qwen 3.8 27B.

The Qwen 3.8 family, one line each:

  • Qwen 3.8 Max (this post): 2.4T flagship (MoE), 80/100 on the only independent benchmark against Kimi K3's 83; access via Token Plan from $6 to $68.
  • Qwen 3.8 Max open weights: what changes with the weights released and what you can (and can't) run.
  • Qwen 3.8 Flash Next 125B-A6B: the fast, cheap sibling, 125B with 6B active.
  • Qwen 3.8 27B: the checkpoint that runs on your machine.

The context: why Qwen 3.8 caught fire

First, the name. If you search for "qwen 3.8" right now, half the results still point to qwen3:8b, that dense 8-billion-parameter model that runs for free on your laptop. That's not it. Qwen 3.8 Max is a different category: 2.4 trillion parameters, three orders of magnitude above. The SERP hasn't separated the two yet, and that's exactly why whoever gets to the right term first educates the market.

Second, the timing. Qwen 3.8 came out two days after Kimi K3, from Moonshot AI, another Chinese model, 2.8 trillion parameters, this one already with open weights. A detail almost nobody mentions: Alibaba owns about 36% of Moonshot. So there's internal portfolio competition going on, with both sides pushing the frontier at the same time.

Third, the backdrop. In just a few months we got Z.AI's GLM 5.2 (which closed in on Claude Opus 4.8), Kimi K3 (which beat Fable 5 on specific benchmarks) and now Qwen 3.8, the biggest of the three. This isn't an isolated launch. It's a wave of Chinese labs aiming at the top, and choosing open weights as their commoditization weapon.

The technical part Alibaba didn't share matters too. How many of the 2.4 trillion parameters are active per token? In a sparse MoE, that's the number that defines the real cost of inference. It wasn't disclosed. "Trust us, it's number 2" is precisely the kind of claim a benchmark table exists to settle, and the table never showed up.

What the community has already tested

This is where the fight gets interesting, because both sides are partly right.

The excited side. On Hacker News, the Qwen 3.8 thread blew up, with the top comment passing 900 points. Devs testing the multimodal image generation were impressed: they asked for an SVG of a pelican and the model, on its own, added a little fish in its beak. A creative detail that wasn't in the prompt. In first impressions on code and full-stack work, some people swear it easily beats Sonnet for day-to-day use.

The skeptical side. The same thread turned into a geopolitical debate and, once the dust settled, the technical criticism left standing was harsh: Qwen has a reputation as a "benchmark specialist," a model that shines on tests and slips in production, especially at following instructions. And the most serious independent test to show up so far doesn't confirm the hype. On a software architecture task with a 60-minute limit, trilogyai compared Qwen 3.8 Max against Kimi K3:

Metric Qwen 3.8 Max Kimi K3
Final score (with penalties) 80/100 83/100
Tool calls 44 53
Repository citations 354 274
Gateway failures 0 2 commands denied

Notice: Qwen was more economical with tool calls and cited the repository more (a sign that it read more context before acting), but it lost to Kimi K3 on the final score after the penalties for factual errors. In other words: good, solid, but not the "crushes everything" model the WAIC stage sold. And Kimi K3, remember, is the model from the house next door that Alibaba itself invests in.

The honest read of the first 24 hours: Qwen 3.8 Max is competent and probably better than 3.7 at code and office workflows. But "second in the world" remains a vendor claim until Artificial Analysis and LMArena run the model.

▪ Clã Beer and Code

Do not just follow the news — master it. Hands-on AI Engineering, live, every week, in the largest community in Brazil.

Join the Clã

How to access it today (and why "open source" comes with an asterisk)

Until the weights land, access is paid and goes through Alibaba's channels. The model shows up as Qwen3.8-Max-Preview.

  • Token Plan: a credit-based subscription, from Lite at $6 (2,500 credits/week) to Pro at $68 (40,000 credits/week). There's no published per-token price, so cost forecasting is left in the dark.
  • Qoder and QoderWork: Alibaba's own agent/IDE tools.
  • Preview at 10% of the standard price, for a limited time.

The API is compatible with the OpenAI and Anthropic protocols. In practice, you swap the base_url and the key, and most SDKs work:

from openai import OpenAI

client = OpenAI(
    base_url="https://dashscope.aliyuncs.com/compatible-mode/v1",
    api_key="YOUR_ALIBABA_KEY",
)

resp = client.chat.completions.create(
    model="qwen3.8-max-preview",
    messages=[
        {"role": "user", "content": "Refactor this endpoint to use caching and explain the trade-off."}
    ],
)
print(resp.choices[0].message.content)

About the "open source Fable" from the headlines: for now it's just a promise. Alibaba said the weights will open "soon," but gave no date and no license. And even when they do open, there's a detail that kills self-hosting for most people: the weights should exceed 1.2 TB. A typical GPU holds 141 GB. So "open" won't turn into "runs on my machine" anytime soon. It'll turn into "runs for whoever has a cluster, or a provider that spins it up for you."

Limitations and things to watch

  • Zero third-party benchmarks. Every performance number today comes from Alibaba. It's not a lie, but it's marketing until independently proven.
  • Active parameters not disclosed. Without that, you can't really estimate inference cost in a MoE.
  • It's a preview. The model card and license don't exist yet. Don't build a critical product on top of something that could change its behavior or its price tomorrow.
  • Instructions in production. The Qwen lineage has a reputation for slipping on instruction-following under load. If you're going to test it, test it on your real use case, with your prompts, not on the pretty benchmark.
  • It's not qwen3:8b. Again: the local 8B model has nothing to do with this 2.4T one. Don't mix them up when you go to download.

Quick FAQ

Is Qwen 3.8 Max open source yet? No. Today it's a proprietary preview via Token Plan, Qoder and QoderWork. Alibaba promised open weights, but with no set date or license.

Can I run it locally? In theory, once it opens. In practice, not for almost anyone: the weights exceed 1.2 TB. That's cluster territory, not laptop territory.

Is it better than Claude? Alibaba says it trails only Fable 5. There's no independent benchmark to confirm that. The only serious test that showed up put it slightly behind Kimi K3 on an architecture task. Wait for Artificial Analysis and LMArena before believing the ranking.

How do I access it now? Via Token Plan (credit-based subscription) or through the Qoder tools. The API is OpenAI/Anthropic compatible: swap base_url and key. Beyond that, you wait for it to land on OpenRouter or for the weights to open.

Conclusion

Qwen 3.8 Max is real, it's big and it's competent. But the distance between "strong model from China" and "second best in the world" is exactly the size of the benchmarks that don't exist yet. Did it get to the term first? It did. Did it prove the title? Not yet.

The pattern here applies to any launch: it's not the parameter count that decides, it's the model running on your problem, with your data, under your evaluation. Test fast, measure with method and don't buy stage hype. That's AI engineering, not rooting for a team. It's the kind of reading we do together in the Clã Beer and Code, the largest AI engineering community in Brazil, where people put new models into production before the SERP figures out what happened.

The next chapter is predictable: when the weights open and Artificial Analysis publishes the table, we'll know whether it was an open source Fable or just the biggest model with the smallest benchmark. Until then, test it yourself, and be suspicious of anyone who has already called the final score.

Lucas Souza
Written by
Lucas Souza

{AI Engineer} — apaixonado por Laravel, arquitetura de software e construir produtos com impacto. Compartilho aqui tutoriais, descobertas e reflexões sobre o dia a dia de engenharia.

▪ Clã Beer and Code

There is no shortage of content. What is missing is someone to untangle it: what matters now is how to implement it the right way. In the Clã you get that live, every week, with people who have already filtered out the noise.

Join the Clã
Meet the Clã Beer and Code
playing