#Modelos Open Source
PewDiePie says OpenAI banned his account twice for distillation while he was training Ajax, a local Qwen 3.5 9B with refusals removed by Heretic. What model distillation is, how OpenAI detects it (the report on the Moonshot case came out two days earlier), what's inside Ajax, and where the line sits between legitimate synthetic data and a ban on your account.
Jev costs $0.042 per million tokens and is a closed API. Laya is Apache 2.0, runs offline on a GPU from 2018, and measures 7.8x faster at P50. A comparison using the numbers each side published, plus the prior-art fight that blew up on Hacker News two days after the launch.
DeepSeek released V4.1 Flash and the timeline cropped out the good row of the benchmark. We compare the model with Opus 5, GPT-5.6 Sol, Kimi K3, GLM-5.3, and GPT-6 Astra on price and performance, show where it actually leads, where it drops 20 points, and the number buried in the model card: the scaffold changes the result forty times more than swapping the model.
Alibaba announced Qwen3.8-Flash-Next: 125B total with only 6B active per token, plus 51B in N-gram embeddings and a redesigned sparse attention. What Qwen has confirmed, what's still community estimate, how much memory it really needs, and why the architecture is being published ahead of Qwen 4. No official benchmark has come out so far.
Ox Alpha was GLM-5.3-Flash. Five days before any announcement, tokenizer fingerprinting was already pointing to Zhipu: 95 out of 95 against the GLM-5 vocabulary. Now Z.ai has confirmed it, published the weights on Hugging Face under an MIT license and revealed the architecture: 320B total with 18B active, 1M context, $0.075 per million. The 80% benchmark is still what it always was: a sample of ten tasks. On the full set, 63%.
Alibaba shipped Qwen 3.8 27B and someone opened the diff against 3.6: 59 of 59 graph nodes map one-to-one, and the only differing field is metadata. Same architecture, DeepSWE tripling from 13.3 to 42.2. Here are the real benchmarks (and what the vendor table leaves out), the VRAM math the press oversimplified, the 64KB-per-token KV cache, the Jinja template bug that kills tool calls on day 1, and the difference between the dense 27B and the 2.4T 3.8 Max. With the counterpoint nobody made.
DeepSeek published DeepSeek-V4-Pro-0813 on its official pricing page, and the official V4 Pro release has finally dropped the preview label. The reported numbers, what's fact and what still has no public document, the comparison with Flash 0731 and Opus 4.8, and the notice on the pricing page itself that DeepSeek is going to raise prices soon.
Cactus Compute shipped Needle 2: 45M parameters, a 14 MB binary, a session in 28 MB of RAM, 500 tokens/s on a Raspberry Pi 5 and 70 MFLOPs per token versus 460 for LFM2.5 230M. The engineering is real: CQ2-bit quantization applied from pre-training onward, a byte-level grammar that locks the output to valid function calls. But the Show HN turned into a public failure lab: typing "HN" fires lock_door with confidence 0, "warmer" becomes mode cool, and the ESP32 demo that went viral was Needle 1.
The press says Muse Glimmer 30B requires a 5090. r/LocalLLaMA is posting screenshots of it running on a used 3090 from 2020. Both are right, and the explanation is in the VRAM budget: 17 GB of weights, 1.7 GB of KV cache, and an attention architecture designed to fit. Here's the math line by line, the tokens-per-second estimate on a 3090 with the work shown, and the verdict on when 24 GB is enough and when it isn't.
Alibaba is publishing the weights of Qwen3.8-Max, 2.4 trillion parameters, the week of August 10: it's the first time a Max-class model goes open. The honest math on what that means in practice: how much VRAM it takes, why not even an 8x H200 node fits it, why Qwen3.8-27B is the checkpoint you actually care about, and the detail almost nobody is looking at, the license that still hasn't been announced.
DeepSeek republished the V4-Flash weights on July 31 without changing the model name in the API: anyone calling deepseek-v4-flash woke up running a different model, with no changelog. The real 0731 numbers (82.7 on Terminal Bench, but 79% in the independent measurement), where it beats GPT-5.6 Luna and where it loses, the 169 GB to run it locally and what to do if your agent points at a model name that became a moving target.