~/beer-and-code
▪ next event Workshop: Jev na Prática para Devs · 14 Oct · 19h — duração de 2h a 3h, ao vivo via Google Meet save your seat ›
~ / news / kimi-k3-benchmarks-vs-claude $
News

Kimi K3 Pricing: $3/$15 per Million Tokens, Is It Free? Benchmarks vs Claude

LS Lucas Souza · · 10 min read
Kimi K3 Pricing: $3/$15 per Million Tokens, Is It Free? Benchmarks vs Claude

Every quarter the same scene plays out: a new Chinese model drops, the timeline declares closed models dead, and two weeks later nobody remembers its name. GLM, DeepSeek, MiniMax, Qwen. We all know the script.

Except this time the numbers didn't come from marketing. Moonshot AI launched Kimi K3 on July 16 and it debuted in 3rd place on the Artificial Analysis index, above Claude Opus 4.8, and in 1st place on the Frontend Code Arena, above even Fable 5. With open weights promised for July 27 and Sonnet-level pricing.

This post separates the verifiable benchmarks from the timeline hype: the real numbers, the honest pricing math (spoiler: the "5x cheaper" claim going around doesn't add up), what the community is still skeptical about, and how to test it today. And since the news is less than 48 hours old, this is a living post: we'll update it as independent tests come out.

TL;DR

  • What it is: Kimi K3, an LLM from Moonshot AI. A 2.8 trillion parameter MoE (16 of 896 experts active), 1 million token context, natively multimodal. The largest open-weight model ever announced.
  • Benchmarks: 57 on the Artificial Analysis Intelligence Index (Opus 4.8: 56) and #1 on the Frontend Code Arena with 1679 points, passing Fable 5.
  • Cost/access: $3 input / $15 output per million tokens (cache hit: $0.30). Already available through the Moonshot API and on OpenRouter as moonshotai/kimi-k3.
  • Weights: full release promised by July 27. Until then, "open" is a promise, not a fact.

Update (July 27): the weights landed a day earlier than promised. The modified MIT license, the native MXFP4 format, and the real math of running it locally (1.4 TB of VRAM, break-even at ~6,200 tokens/s against the API) are in Kimi K3 released its open weights: what does it run on?.

What Moonshot actually launched

K3 isn't just "one more big model." The official announcement includes architecture decisions that explain how a 2.8T model can be served at Sonnet pricing:

  • Aggressive MoE: of the 2.8 trillion parameters, only 16 of 896 experts activate per token. You pay (in compute) for a fraction of the total size.
  • Kimi Delta Attention (KDA): a hybrid linear attention that, according to Moonshot, decodes up to 6.3x faster at 1M tokens of context.
  • Native quantization: weights in MXFP4 and activations in MXFP8, with quantization-aware training starting at SFT. The model is born optimized for cheap inference; it isn't quantized after the fact.

Concept, application, impact: sparse architecture + linear attention + native quantization is what turns "a giant model that exists" into "a giant model you can sell at $3 per million tokens." It shipped in two variants: K3 Max (chat and agents) and K3 Swarm Max (parallel processing at scale), according to SiliconANGLE.

The market felt it: AI and semiconductor stocks fell on Friday as investors relived the "DeepSeek moment" of 2025.

The benchmarks: where Kimi K3 actually beats Opus 4.8

Let's stick to what's verifiable, with sources:

  • Artificial Analysis Intelligence Index: K3 scores 57, against 56 for Claude Opus 4.8 (adaptive reasoning, max effort). That puts K3 in 3rd among the models in the ranking, behind only Claude Fable 5 (59.9 in the Max version) and GPT-5.6 Sol. It's the first time an open-weight model has passed a current-generation Opus on the aggregate index. Head-to-head comparison here.
  • Frontend Code Arena: #1 with 1679 points, passing Fable 5, and a 17-position jump over K2.6, which sat at #18. K3 took 1st in 6 of the 7 frontend domains (Brand & Marketing, Reference-Based Design, Data & Analytics, etc.), losing only Gaming to Fable 5.
  • Moonshot's own benchmarks: 90.4 on BrowseComp and 67.3 on DeepSWE, according to the official announcement. Vendor numbers: treat them as claims until someone replicates them.

Notice the difference in kind between these sources. Artificial Analysis and Arena are third-party measurements; Moonshot's table is launch material. The first two can carry a headline; the third can carry a hypothesis. If you want to build that reading muscle, we broke the method down in how to read benchmarks in practice.

Kimi K3 pricing (and the "5x cheaper" that doesn't add up)

Here's the part nobody checks. The timeline is saying K3 delivers "Opus level for 5x less." Let's do the math with the official prices:

Model Input ($/Mtok) Output ($/Mtok) Weights
Kimi K3 3.00 (0.30 with cache hit) 15.00 Open (by July 27)
Claude Opus 4.8 5.00 25.00 Closed
Claude Fable 5 10.00 50.00 Closed
GLM 5.2 1.40 4.40 Open (MIT)
Kimi K2.6 0.95 4.00 Open

The real math: K3 costs 1.7x less than Opus 4.8 and 3.3x less than Fable 5. No honest comparison gets you to 5x. Is it cheaper than the closed top tier? Yes. Is it the bargain the timeline is selling? No. Even the pricing hype came benchmaxxed.

And there's the opposite move that almost nobody mentioned: K3 is 3x more expensive than K2.6 itself ($0.95/$4.00) and twice the price of GLM 5.2. As The Decoder pointed out, K3 signals the end of the ultra-cheap Chinese AI era: when the model closes in on the frontier, the price closes in too. It makes sense. Serving 2.8T parameters, even sparse ones, isn't free.

For the dev running a coding agent all day, the practical reference point: $3/$15 is exactly the price of Claude Sonnet 5. In other words, Moonshot's bet is to deliver Opus results while charging Sonnet prices.

▪ Clã Beer and Code

Do not just follow the news — master it. Hands-on AI Engineering, live, every week, in the largest community in Brazil.

Join the Clã

"Benchmaxxed"? The skepticism that remains

Every Chinese launch goes through the same ritual: spectacular numbers on day 1, and the inevitable question: is the model good, or was it trained to look good on benchmarks? The community's term is benchmaxxed, and the distrust has history behind it: we've seen models fly on a leaderboard and trip over the first real repository.

What we have in independent testing so far is thin, and it points to nuance. Simon Willison ran K3 on the classic pelican-riding-a-bicycle SVG test: the result came out good, but the model burned 13,241 reasoning tokens on a simple task, $0.25 on a 95-token prompt. Willison himself adds the caveat that matters: a drawing benchmark doesn't measure what defines a model in 2026, which is agentic tool calling and reliability over long sessions. And that's exactly where there's still no independent evaluation of K3.

Here's the honest summary: the third-party numbers (Artificial Analysis, Arena) are real and impressive. What's missing is the test no benchmark captures: turning the model loose in an agent harness, on a real codebase, for hours. We ran exactly that kind of test in the Opus 4.8 vs Minimax M3 vs local Qwen 3 comparison, and the lesson from there applies here: a leaderboard number tells you who to invite to the tryout, not who to hire.

[UPDATE: this block will be edited as independent tests of K3's agentic use come out.]

The backdrop: open source became state policy

The timing of the launch is no accident. Days earlier, Xi Jinping opened WAIC in Shanghai, his first appearance at the event, saying that AI "should not be a solo show by a single country, but a symphony of international cooperation" and making open source an explicit directive of China's AI strategy.

Translating the geopolitics into engineering: Chinese labs don't open their weights out of altruism, they open them because it's the competitive weapon available against the closed American trio. And that produces a great side effect for people who build: every quarter, the best model you can download and run on your own infra gets better. That's how it went with GLM 5.2 closing in on Opus at code, and that's how it's going now with K3 passing Opus on the aggregate index.

How to test Kimi K3 today

Without waiting for the weights:

  • API/OpenRouter: the model is already on OpenRouter as moonshotai/kimi-k3, at $3/$15, served by Moonshot itself. Plug in the key, point your harness at it, test.
  • App and web: available at kimi.com and in Moonshot's apps.
  • Self-host: only after July 27, when the weights come out. And the usual warning applies: 2.8T parameters, even in MXFP4, is serious infra. Expect requirements in the range of multiple GPU nodes, not "runs on a 4090."

One caution we repeat for every Chinese model: using the API means your data travels to Moonshot's infra. For proprietary code or customer data, read the terms first, or wait for the weights and run it at home, which is precisely the superpower of open-weight.

Quick FAQ

Is Kimi K3 free? The chat at kimi.com has a free tier. The API is paid: $3 per million input tokens and $15 per million output tokens. The weights will be opened by July 27, at which point running it "for free" becomes a question of how much your GPU costs.

Does Kimi K3 work well in Portuguese? Moonshot hasn't published a Portuguese-specific benchmark. The previous generation (K2) was already usable in pt-BR, and models this size tend to do well in Romance languages. Our Portuguese test goes into the next update of this post.

Can I use it commercially? Through the API, yes, under the Moonshot/OpenRouter terms. Through the weights, it depends on the license that ships on the 27th. The expectation is Modified MIT, as with K2, which allows commercial use. Confirm the published license before building a product on it.

Is Kimi K3 better than Claude? On the aggregate indexes it passes Opus 4.8 and sits behind Fable 5 and GPT-5.6 Sol. On frontend, it leads the Arena. But "better" for people who build is whatever survives in your use case. Run the candidate in your pipeline with your data before migrating anything.

Conclusion: hype or turning point?

Both, and that's why the news matters. The turning point is real: an open-weights model debuted above Opus 4.8 on an independent index and at the top of a code arena. That had never happened before. The hype is real too: the "5x cheaper" claim doesn't survive a simple division, Moonshot's own benchmarks haven't been replicated yet, and the test that actually settles it, an agent running for hours on production code, nobody has published yet.

Our recommendation is the usual one: don't migrate anything because of a headline. Put moonshotai/kimi-k3 in your harness, run your evals, compare it with what you use today, the way we did when Fable 5 came back and the math on which model to use changed again. Testing a new model in a real pipeline, by the way, is what we do every week, live, in the Clã Beer and Code, the largest AI engineering community in Brazil, with people applying this in real products, not just reading leaderboards.

This post will be updated as the weights (July 27), the final license, and the first independent agentic tests come out. Check back.

Lucas Souza
Written by
Lucas Souza

{AI Engineer} — apaixonado por Laravel, arquitetura de software e construir produtos com impacto. Compartilho aqui tutoriais, descobertas e reflexões sobre o dia a dia de engenharia.

▪ Clã Beer and Code

There is no shortage of content. What is missing is someone to untangle it: what matters now is how to implement it the right way. In the Clã you get that live, every week, with people who have already filtered out the noise.

Join the Clã
Meet the Clã Beer and Code
playing