~/beer-and-code
▪ next event Workshop: Jev na Prática para Devs · 14 Oct · 19h — duração de 2h a 3h, ao vivo via Google Meet save your seat ›
~ / news / kimi-k3-open-weights $
News

Kimi K3 Open Weights Are Out: The World's Largest Model Is Free, but What Does It Run On?

LS Lucas Souza · · 14 min read
Kimi K3 Open Weights Are Out: The World's Largest Model Is Free, but What Does It Run On?

Moonshot AI put up a countdown promising the Kimi K3 open weights for July 27. It shipped on the 26th, around 7:30 PM New York time, a day ahead of its own deadline. It's not every day a lab delivers early.

The file is there: 2.8 trillion parameters, the largest open-weights model ever published, free to download. And that's where the part the hype leaves out begins. Because "free" and "runs on your machine" are two completely different claims, and the gap between them, in this case, is roughly 1.4 terabytes of VRAM.

In this post: what exactly Moonshot released (license, format, where to download), the hardware a MoE at this scale actually demands, the honest math of running it locally versus paying the API at $3 / $15 per million tokens, and what US sanctions can reach, or can't, when you download weights from Brazil.

TL;DR

  • What it is: Kimi K3, a MoE with 2.8T total parameters and 104B activated, a 1,048,576-token context window, native vision. Weights published on July 26, 2026.
  • License: Modified MIT. Internal and commercial use are allowed; operating K3 as Model-as-a-Service with revenue above $20 million over 12 months requires a separate agreement with Moonshot.
  • Format: safetensors with native MXFP4 quantization (weights) and MXFP8 on activations, applied through quantization-aware training. This is not a quantization done after the fact by a third party.
  • Real cost of running it: ~1.4 TB just to load the weights. It doesn't fit in any 8-GPU node of the H200 generation, and not in a DGX B200 either.
  • API: $3 per million input tokens (cache miss), $0.30 on cache hit, $15 on output. No long-context surcharge.

What Moonshot actually released in the Kimi K3 open weights

Let's go to what's on the Hugging Face card, because that's where the difference between news and rumor lives.

Architecture. 2.8 trillion total parameters, 104 billion activated per token. There are 896 experts, 2 of them shared, and the router activates 16 per token. 93 layers, of which 69 use Kimi Delta Attention (the house hybrid linear attention) and 24 use Gated MLA. A MoonViT-V2 vision encoder with 401M parameters. A 1,048,576-token context window.

Format. Safetensors with native MXFP4 on the weights and MXFP8 on the activations. And there's an engineering detail here that matters: the quantization was applied with quantization-aware training, not bolted on afterward to a finished FP16 model. In other words, the MXFP4 is not a degraded version of the "real model". It is the real model. You are not downloading the GGUF somebody made over the weekend.

License. It's a modified MIT, and the modification is what you need to read. Two clauses take the "MIT" out of the conversation if your product grows:

  1. If you (or an affiliate) operate a Model-as-a-Service business and aggregate revenue exceeds $20 million in any 12 consecutive months, you need to sign a separate agreement with Moonshot before using it.
  2. A commercial product with more than 100 million monthly active users or more than $20 million in monthly revenue has to display "Kimi K3" prominently in the interface.

The limits do not apply to internal use, defined as any use that does not make the model, its outputs or its capabilities available to third parties. Translated to your case: if you're going to run K3 inside the company, in an internal pipeline, in a product of yours that doesn't resell inference, none of this touches you. If your business is selling tokens, read the license with your legal team.

To size up what this means competitively, it's worth reading Nathan Lambert, who called the launch a "watershed moment" and estimates that the gap between open and closed models dropped from 6 to 9 months to something between 3 and 5. His line that sums up the discomfort in Washington and San Francisco: "Moonshot AI is going toe to toe with Anthropic and OpenAI with far, far fewer resources".

But published weights are not a product. Between downloading a safetensors file and having inference serving requests at acceptable latency there is an entire layer of engineering, and that layer is exactly what separates people who use AI from people who build with AI. That's the subject of the Clã Beer and Code, the largest AI engineering community in Brazil: it's paid, it's a subscription, and it's where this math shows up with numbers on the table instead of an enthusiasm thread.

If you want the capability angle, meaning how much K3 really delivers against Claude and GPT on benchmarks, we already broke that down when the model launched via API. This post is about the other half of the story: what it costs to get this thing running.

Kimi K3 requirements: the 1.4 TB of VRAM that doesn't fit in the headline

Here is the number that decides everything: the MXFP4 weights take up about 1.4 TB. In FP16 they would be somewhere near 5.6 TB. The native quantization already gave you a 4x discount, and you're still left with 1.4 TB to load before processing a single token.

Do the math with the hardware that exists:

Configuration Aggregate VRAM Fits?
1× RTX 5090 (32 GB) 0.032 TB No, not even close
8× H100 (80 GB) 0.64 TB No
8× H200 (141 GB) 1.128 TB No
8× B200 (180 GB) 1.44 TB Technically yes, in practice no
8× B300 / MI355X (288 GB) 2.304 TB Yes
GB300 NVL72 ~20.7 TB Yes, with room to spare

Look at the H200 row. There's a lot of talk going around that "an 8× H200 node does the job if you accept MXFP4". It doesn't: 8 × 141 GB is 1.128 TB, and the weights need 1.4 TB. You're short before you start.

The B200 row is subtler and more interesting. 8 × 180 GB is 1.44 TB. The weights fit, with about 40 GB left over across the whole node. Then you remember you still need room for KV cache, activations and runtime overhead, and that the model advertises 1M of context. That's why SemiAnalysis was blunt: K3 does not fit in a DGX B200 even in FP4, and the path is GB300 NVL72, B300 or MI355X, where each GPU has 288 GB.

Kimi Delta Attention helps a lot here. Hybrid linear attention uses far less KV cache than full attention at the same length, and that's why 1M of context is viable at all. It helps. It doesn't work miracles in 40 GB.

Can you go with two 8× H200 nodes (16 GPUs, 2.256 TB) and distribute? You can. But then you hit the other wall, which SemiAnalysis itself points out: the B200 has 400 Gbit/s of bandwidth between nodes, versus an NVL72 with roughly 18x more inter-node bandwidth. In a MoE with 896 experts, token routing crossing the network at every layer is not a detail. It's the bottleneck. A sparse model punishes bad interconnect in a way a dense model doesn't.

And the runtime? vLLM, SGLang and TensorRT-LLM are the supported paths, with the caveat that routing across 896 experts requires MoE-aware scheduling, so an outdated engine may need a patch. It's not docker run and done.

Running Kimi K3 locally vs the API at $3 / $15: where the break-even sits

Now the math that matters to whoever decides the budget.

The rental side. Market prices in July 2026: H200 runs between $3.72 and $10.60 per GPU-hour, with a median near $3.95. CoreWeave lists the 8-GPU HGX H200 node at $50.44/hour on demand, and about $20.93/hour on spot. B200 sits in the $6 to $8.60 per GPU-hour range. GB200 NVL72 goes from $10.50 to $27 per GPU-hour.

Take the minimum configuration that satisfies the memory math with two H200 nodes:

  • On demand: 2 × $50.44 = $100.88/hour → $2,421/day → ~$72,600/month running 24/7.
  • Spot: 2 × $20.93 = $41.86/hour → ~$30,100/month, with the preemption risk that comes with spot.

The API side. $3 per million input, $15 per million output, $0.30 on cache hit. No long-context surcharge, which, on a 1M model, is no small thing.

Let's use a realistic blend for an agentic workload, which is input-heavy: for every 1M input tokens, about 100k output. Cost per 1M input-equivalent = $3 + (0.1 × $15) = $4.50.

Now the break-even:

$72,600 / $4.50 per 1M tokens = ~16.1 billion tokens/month
16.1e9 tokens / (30 days × 86,400 s) = ~6,200 tokens/second, 24 hours a day, 30 days

Six thousand two hundred tokens per second. Sustained. Nonstop. That's the volume you need to maintain for on-demand rental to match the API. In the spot version, the number drops to ~2,600 sustained tokens/s, which is still a company-sized workload, not a team-sized one.

And there's a detail that makes the local side look even worse: the cache hit at $0.30/M. If your workload has a repeated prefix (a large system prompt, a codebase, a fixed document, which is exactly the agent profile), the API gets 10x cheaper on that slice. The break-even climbs with it.

The uncomfortable conclusion is simple: running Kimi K3 locally is not a cost decision. At this scale, the API wins by a landslide for almost everyone. Local wins for other reasons, and they are legitimate: data sovereignty, a regulatory requirement that forbids inference from leaving the perimeter, fine-tuning and continued pre-training on the weights, or research. If your argument for self-hosting is "I'll save money", redo the spreadsheet.

That's different from what happens one step down. A smaller model on your own hardware is another conversation. We showed DeepSeek V4 Flash running 1M of context on an RTX 5090, with that case's asterisks, and compared paying a lot, paying a little and running local for free while building the same product three times. There, "running locally" is a phrase that means something. At 2.8T, "local" became a synonym for "datacenter".

For anyone who wants to test without building a cluster: Together AI and Modal announced day-zero hosting, alongside Moonshot's own API.

▪ Clã Beer and Code

Do not just follow the news — master it. Hands-on AI Engineering, live, every week, in the largest community in Brazil.

Join the Clã

What sanctions can reach when you download weights from Brazil

A question that has shown up in every thread since the 21st, when Treasury Secretary Scott Bessent threatened to sanction Chinese labs over alleged intellectual property theft: does downloading this put me at risk?

Short answer: today, no. There is no formal published sanction against Chinese open-weights models. There is a public threat, which is something else. And even if one comes, there's a property of the format that changes everything: published weights are irreversible. The license is an irrevocable grant, the file is on your disk, and there is no kill switch. A US sanction doesn't erase safetensors from your storage, and it doesn't turn local inference into something illegal in Brazil.

The real risk is indirect, and it's worth mapping:

  • Distribution. Hugging Face is an American company. A takedown or geo-block of the repository is the plausible scenario. If K3 is on your roadmap, download and version it now. Storage is cheap; re-downloading after a takedown doesn't exist.
  • Providers. Together, Modal, Fireworks, OpenRouter: American companies comply with American sanctions. Your access through a third-party API is the most fragile point of failure in the chain.
  • Customer compliance. You may be a Brazilian-registered company (you have a CNPJ, the Brazilian corporate tax ID), but if you sell to a multinational with American capital, the customer's legal team can veto a sanctioned model in the supplier chain. The contract reaches where the sanction doesn't.

We broke down the full scenario (what's fact, what's rumor and the minimum playbook) in the post on sanctions against Chinese open source AI. The operational summary still holds: versioned weights, abstracted provider, your own evals, a fallback defined ahead of time with a cool head.

And don't confuse the two things: the $20 million limit in Moonshot's license is a contractual clause, not a geopolitical sanction. One you solve with a lawyer and an agreement; the other doesn't depend on you.

Limitations and things to watch

Quantization is not a ladder. Since MXFP4 is native and trained, the comfortable path of "I'll take the FP16 and quantize further" doesn't exist. You're already at the source format. More aggressive third-party quantizations will show up, but they start from a model that was already optimized for 4 bits, so the margin to cut without breaking things is much smaller than what existed with FP16 models.

1M of context is a ceiling, not a default. Advertising 1,048,576 tokens doesn't mean you'll serve everyone a full window. Every context token takes up memory that competes with the model's weights on the same node. In production, you'll set a much lower limit per request and you'll like the result.

Engine support is still maturing. 896 experts with 2 shared and 16-per-token routing is a new topology. Expect vLLM/SGLang versions with specific fixes in the coming weeks, and don't plan a critical deploy on top of the first release that works.

A benchmark is not your use case. True for every model, doubly true here: third place in a general ranking says nothing about your domain. Without your own eval suite, "K3 is better" is an opinion.

Quick FAQ

Can I run Kimi K3 on my RTX 5090 with some quantization? No. The native MXFP4 weights take ~1.4 TB; a 5090 has 32 GB. Not even aggressive offloading to RAM and disk makes this usable. You'd be trading millisecond latency for minutes. For real local use, look at models from 30B to 200B.

Can I use it commercially? Yes, in most cases. The modified MIT allows internal and commercial use. The restrictions only kick in if you operate Model-as-a-Service with revenue above $20M/12 months, or have a product with 100M+ monthly users (in which case you need to credit "Kimi K3" in the interface). Internal use is explicitly exempt.

Where do I download Kimi K3 and what's the real download size? From the moonshotai/Kimi-K3 repository, in MXFP4 safetensors, in the neighborhood of 1.4 TB. Plan bandwidth and storage before you run git clone, and use hf download with resume, not the browser.

Is it worth building your own infrastructure for this? Only if the reason isn't cost. Per the break-even above, you need ~6,200 tokens/s sustained 24/7 to match the on-demand API. If the reason is data sovereignty, a regulatory requirement or fine-tuning on the weights, then yes, the math changes in nature. It stops being about savings and becomes about capability the API doesn't sell.

Conclusion

The open-weights Kimi K3 is the largest model ever freely published, it came out a day earlier than promised, with a license most teams can use without talking to a lawyer, and in a format that isn't a quantization hack. That is real and it's big.

What isn't real is the idea that "open" solved the access problem. It solved the permission problem. The capacity problem is exactly where it was: 1.4 TB of VRAM, datacenter interconnect and a break-even that only closes at company scale. The model became free; the infrastructure to run it stayed exactly the same.

And that's the provocation this launch leaves behind. When a frontier model is free to download but requires a $3 million rack to serve, the market's bottleneck stops being "who has the best model" and becomes "who can operate the model that already belongs to everyone". That's not a research problem. It's an engineering problem, and nobody is going to publish that part on Hugging Face for you.

Lucas Souza
Written by
Lucas Souza

{AI Engineer} — apaixonado por Laravel, arquitetura de software e construir produtos com impacto. Compartilho aqui tutoriais, descobertas e reflexões sobre o dia a dia de engenharia.

▪ Clã Beer and Code

There is no shortage of content. What is missing is someone to untangle it: what matters now is how to implement it the right way. In the Clã you get that live, every week, with people who have already filtered out the noise.

Join the Clã
Meet the Clã Beer and Code
playing