#Hardware
It's not a new model. It's the same GPT-5.6 Sol running on a chip the size of a dinner plate: 750 tokens/s, 44 GB of on-chip SRAM and zero published pricing. OpenAI claims 14x against its own Sol Standard; Cerebras claims 11x against Fable 5 (with no head-to-head test). Here we separate what's verifiable from vendor marketing, explain why Sol, Ultra and Ultrafast are three different things, and walk through the math that decides whether latency turns into money in your agent.
Alibaba shipped Qwen 3.8 27B and someone opened the diff against 3.6: 59 of 59 graph nodes map one-to-one, and the only differing field is metadata. Same architecture, DeepSWE tripling from 13.3 to 42.2. Here are the real benchmarks (and what the vendor table leaves out), the VRAM math the press oversimplified, the 64KB-per-token KV cache, the Jinja template bug that kills tool calls on day 1, and the difference between the dense 27B and the 2.4T 3.8 Max. With the counterpoint nobody made.
Cactus Compute shipped Needle 2: 45M parameters, a 14 MB binary, a session in 28 MB of RAM, 500 tokens/s on a Raspberry Pi 5 and 70 MFLOPs per token versus 460 for LFM2.5 230M. The engineering is real: CQ2-bit quantization applied from pre-training onward, a byte-level grammar that locks the output to valid function calls. But the Show HN turned into a public failure lab: typing "HN" fires lock_door with confidence 0, "warmer" becomes mode cool, and the ESP32 demo that went viral was Needle 1.
The press says Muse Glimmer 30B requires a 5090. r/LocalLLaMA is posting screenshots of it running on a used 3090 from 2020. Both are right, and the explanation is in the VRAM budget: 17 GB of weights, 1.7 GB of KV cache, and an attention architecture designed to fit. Here's the math line by line, the tokens-per-second estimate on a 3090 with the work shown, and the verdict on when 24 GB is enough and when it isn't.