~/beer-and-code
▪ next event Workshop: Jev na Prática para Devs · 14 Oct · 19h — duração de 2h a 3h, ao vivo via Google Meet save your seat ›
~ / tag / #china $ grep

#China

7 posts
01 #llm · #seguranca-ia
Anthropic accuses Moonshot of serving Claude instead of Kimi: what the report proves (and what it doesn't)

Anthropic accuses Moonshot of relaying Kimi customer requests to Claude Opus and showing the answer as if it came from Kimi: almost 300,000 requests in ten days, 5,380 fraudulent accounts and 23 million distillation exchanges. We separate what the 154-page report says, what it doesn't prove, what Moonshot has (not) answered and what changes for anyone running Kimi K3 in production.

06 Oct · 15 min ›
02 #ia · #llm
Ox Alpha Was GLM-5.3-Flash: Z.ai Confirmed It, Opened the Weights Under MIT — and the 80% Benchmark Is Still Fake

Ox Alpha was GLM-5.3-Flash. Five days before any announcement, tokenizer fingerprinting was already pointing to Zhipu: 95 out of 95 against the GLM-5 vocabulary. Now Z.ai has confirmed it, published the weights on Hugging Face under an MIT license and revealed the architecture: 320B total with 18B active, 1M context, $0.075 per million. The 80% benchmark is still what it always was: a sample of ten tasks. On the full set, 63%.

25 Aug · 13 min ›
03 #llm · #custos
DeepSeek V4 Pro 0813: 87.9 on Terminal Bench Beats Opus 4.8, and Your Bill Is Going Up

DeepSeek published DeepSeek-V4-Pro-0813 on its official pricing page, and the official V4 Pro release has finally dropped the preview label. The reported numbers, what's fact and what still has no public document, the comparison with Flash 0731 and Opus 4.8, and the notice on the pricing page itself that DeepSeek is going to raise prices soon.

12 Aug · 11 min ›
04 #llm · #custos
DeepSeek V4 Flash 0731: New Weights, 82.7 on Terminal Bench and the Bill That Hurts OpenAI

DeepSeek republished the V4-Flash weights on July 31 without changing the model name in the API: anyone calling deepseek-v4-flash woke up running a different model, with no changelog. The real 0731 numbers (82.7 on Terminal Bench, but 79% in the independent measurement), where it beats GPT-5.6 Luna and where it loses, the 169 GB to run it locally and what to do if your agent points at a model name that became a moving target.

31 Jul · 12 min ›
05 #llm · #modelos-open-source
Kimi K3 Open Weights Are Out: The World's Largest Model Is Free, but What Does It Run On?

Moonshot AI published the Kimi K3 open weights a day ahead of schedule: 2.8 trillion parameters, a modified MIT license and native MXFP4 format. The math the hype skips: 1.4 TB of VRAM just to load the weights, an 8-GPU node that doesn't add up, and a break-even of 6,200 tokens per second to match the API at $3 / $15. Plus what US sanctions can actually reach when you download weights from Brazil.

27 Jul · 14 min ›
06 #ia · #llm
Qwen 3.8 Max (2.4T): Benchmark vs Kimi K3, Pricing and Open Weights

Alibaba dropped Qwen 3.8 Max, a 2.4-trillion-parameter MoE it claims is the second-best model in the world, behind only Fable 5. With no public benchmark at launch, the only independent test scored it 80/100 against Kimi K3's 83. Here: what it is, what it costs (Token Plan from $6 to $68), how to access it via API, and where each model in the Qwen 3.8 family fits.

20 Jul · 9 min ›
07 #ia · #llm
Kimi K3 Pricing: $3/$15 per Million Tokens, Is It Free? Benchmarks vs Claude

Kimi K3 costs $3 input and $15 output per million tokens on the API (the chat at kimi.com is free), 1.7x cheaper than Claude Opus 4.8, not the "5x" your timeline is claiming. We did the honest pricing math, separated the verifiable benchmarks from the hype, and show where it beats Claude (and where it doesn't).

17 Jul · 10 min ›
Meet the Clã Beer and Code
playing