GPT-5.6 Sol Ultrafast: 750 tokens/s on Cerebras, 11x faster than Fable 5
It's not a new model. It's the same Sol on a different chip
OpenAI did not ship a new model on August 13, 2026.
It shipped the same GPT-5.6 Sol running on a chip the size of a dinner plate. Same intelligence, same context window, same quality. Except the tokens come out at 750 per second.
The name is GPT-5.6 Sol Ultrafast. It's not a model, it's not a reasoning mode — it's an inference tier, running on Cerebras's Wafer-Scale Engine instead of GPUs. It's the kind of launch that separates people who read the announcement from people who read the footnote: headlines from two vendors with different numbers, benchmarks run by the vendor itself, and zero published pricing. Let's separate what's verifiable from the marketing and see where this speed turns into money.
TL;DR
- What it is: an inference tier for GPT-5.6 Sol on Cerebras hardware (CS-3), at up to 750 output tokens per second. Not a new model.
- What changes: only the speed. Intelligence, the ~1M-token context and quality stay the same.
- Access: limited preview, invite-only, form at
/openai-ultrafast-signup. Launch customers: Jane Street, Basis, Rogo and Podium. - Cost: not disclosed. No rate, quota, region, SLA or minimum commitment.
- Primary sources: the Cerebras blog and OpenAI's announcement, "Previewing Ultrafast mode: GPT-5.6 Sol at up to 14X the speed" — a page that returned a 403 and was reconstructed from converging secondary sources.
Sol, Ultra and Ultrafast: three different things with similar names
Sol is the model. $5 per 1M input tokens and $30 per 1M output tokens, with a ~1 million token window. (There's a reseller floating $7/$43 around — ignore it, that's markup.)
Ultra is a reasoning and orchestration mode, not a model: four subagents cooperating in parallel, coordinated by Sol itself. GPT-5.6 Sol Ultra in Codex explains the mechanics and why it burns more tokens per task.
Ultrafast is an inference tier. It changes the hardware underneath, not the brain on top.
In short: ultra changes how much the model thinks; Ultrafast changes how fast it talks. One has a published per-token price. The other still hasn't said what it costs.
And that's where the real decision lives. Buying latency means paying a premium on an input you already use — you don't settle that with a benchmark, you settle it with margin math and architecture design. That conversation, with a CTO, a PM and a SaaS owner looking at the same spreadsheet, is what happens every week in the Clã Beer and Code — because with AI in the mix, the stack stopped being the problem and the decision became a product decision.
GPT-5.6 Sol Ultrafast at 750 tok/s: two headline numbers, two different baselines
OpenAI announced "up to 14X." The baseline is Sol itself in Standard mode, which runs at ~53 tokens/s. Do the math: 53 × 14 ≈ 750.
Cerebras announced "11x faster than Fable 5" and "5x faster than Opus 4.8 in Fast mode." Different baseline: a competitor, not itself.
Both are real and measure different things. But the 11x comes with a caveat almost nobody repeated: Cerebras's literal wording is "compared with output speeds reported by Artificial Analysis" for the Anthropic models. There was no head-to-head test. That doesn't invalidate the number. It changes what the number is.
The benchmarks reinforce the doubt:
- GDP-Val: 5.6x end-to-end, "with no quality loss" — run by Cerebras itself on July 31, 2026.
- HLE: 2,500 questions in 11h11 (Sol Ultrafast) versus 78h27 (Fable 5). Except Sol was measured on July 10 and Fable between July 13 and 15 — different days, unknown load.
And HLE has a structural problem that woadwarrior01, on Hacker News, nailed better than any press release: answering 2,500 independent questions is embarrassingly parallel. Wall-clock time there measures how much capacity you allocated, not how fast the chip is. The honest counterpoint, from rpdillon: "750t/s tells the story".
Rule of thumb: quote the 750 tok/s, take the 11h11 vs 78h27 with a grain of salt. No independent evaluator has measured Ultrafast to date.
Wafer-Scale Engine: 44 GB of SRAM and the bottleneck that moves
On a GPU, the weights live in HBM and get pulled into compute on every token — memory bandwidth is the bottleneck. On the WSE-3 (4 trillion transistors, ~125 petaflops), the weights sit in 44 GB of on-chip SRAM. No round trip. That's where the speed comes from.
Except 44 GB isn't much for long context. philipportner, on HN, did the math Cerebras didn't:
"You'd need hundreds of GB alone for the KV cache of each user. For something like LLama 3 405B you need ~67GB at ~130k tokens. A single CS-3 has 44GB on-chip sram."
Translation: the KV cache of one user with long context already doesn't fit. The model advertises a 1M window; the fast memory has 44 GB. The "up to" in "up to 750 tokens/s" is doing a lot of heavy lifting. Nobody has published the speed × context curve — that's the question you ask before signing a contract.
What keeps the product standing is that the architecture is not "everything on Cerebras." It's hybrid: prefill on AWS Trainium, decode on the CS-3, over Elastic Fabric Adapter — the disaggregated inference announced by AWS and Cerebras on March 13, 2026. The memory-hungry part leaves the wafer. It helps, it doesn't eliminate the problem.
Do not just follow the news — master it. Hands-on AI Engineering, live, every week, in the largest community in Brazil.
Join the ClãLimited preview: who gets in and why there's no price
Access is invite-only. Launch customers: Jane Street, Basis, Rogo and Podium — trading and fintech. Cerebras promises to expand "as capacity grows" and keeps a form at /openai-ultrafast-signup: it's not a waitlist with guaranteed entry, it's a qualification queue. And let's be honest — this product was not designed for the average PHP/Laravel dev. As btown summed it up on HN, there are industries that pay absurd multiples over the API rate for low latency. You are not one of them.
The price is going to hurt because of engineering, not greed. On GPUs, batching spreads fixed cost across many users. On Cerebras, inter-chip bandwidth is ~150 GB/s versus ~2 TB/s for NVLink — in the words of porridgeraisin, you can batch, "but that would just service more users at lower token/s each without any amortization of fixed cost". Every fast user occupies the machine. The only public anchor is the Fast tier: 2x Standard pricing for up to 2.5x the speed. As explainx warns, no per-token Ultrafast number circulating today is reliable.
Context: this is the commercial tip of a contract worth more than $10 billion between OpenAI and Cerebras, from January. ⚠️ The "$20B" figure going around is unconfirmed press reporting.
Where speed turns into money (and where it doesn't)
The math that matters isn't tokens per second. It's time per completed task. And the case where the math works out is an agent in a long loop — 40 sequential calls with 800 output tokens each:
# Conta de guardanapo — só a geração, ignorando prefill e overhead de rede
passos = 40
tokens_por_passo = 800
standard = passos * tokens_por_passo / 53 # ~604 s -> 10 min
ultrafast = passos * tokens_por_passo / 750 # ~43 s
# 10 minutos de espera viram 43 segundos.
# O ganho só existe porque as chamadas são SEQUENCIAIS.
If they were parallel, you'd solve the same thing with GPU concurrency and a smaller bill. Latency only turns into money when the next decision depends on the previous token: an incident while the outage is still happening, interactive financial research, voice AI, multi-step support — the use cases OpenAI named. Jeffrey Wang, an OpenAI researcher, sums up the effect: the task "finishes for me before I even have the opportunity to context-switch".
Now the takeaway the announcement doesn't give you. sixtyj, on HN:
"Compilation time will be a genuine bottleneck for slop coding if this becomes the standard generation rate."
If the model spits out 750 tok/s, composer install, the build and the test suite become the new queue. Generation speed is not delivery speed. Measure how much of your agent cycle is model and how much is pipeline. On most teams, the model stopped being the bottleneck a long time ago.
Limitations and things to watch
Where you get burned:
- No benchmark is independent. GDP-Val and HLE were run by Cerebras, on different dates per model, under the company's verbatim disclaimer: "Performance comparisons are based on third-party benchmarking or internal testing."
- Long context and 750 tok/s probably don't coexist. 44 GB of SRAM against a 1M window, and nobody has published the speed × context curve.
- No price, no SLA, no GA, no public model ID — and no batching economics: unit cost tends to be worse than GPU by design.
- Speed per token is not time per task.
aetherspawnraised the (unmeasured) hypothesis on HN that Sol spends 10–100x fewer output tokens than Fable on the same task. If that holds, the "11x" measures the wrong thing twice — same logic as price per token not being price per task. - Omission flagged, not verified:
ricardobeatcites Mimo v2.5-Pro Ultraspeed, from June, at ~1000 tok/s for less than 1/10 of the cost. No primary source mentions it. - Open gap: nobody has said whether Ultrafast combines with
ultramode. It's the most obvious question of the launch and both primary sources are silent.
FAQ
Is Ultrafast smarter than regular Sol?
No. Same model, same window, same max output — only the delivery speed changes. If your problem is reasoning quality, the path is ultra mode, and there you burn more tokens, not fewer.
How much does it cost? It hasn't been disclosed. No rate, no quota, no SLA. The only public anchor is the Fast tier: 2x the price for up to 2.5x the speed.
Can I use it in ChatGPT or Codex? Nothing has been announced along those lines. The preview is an API preview, invite-only, with trading and fintech at the front of the line.
Is it worth waiting for this for my agent? Only if the loop is long and sequential, and only if you've already measured that generation is the bottleneck. If the cycle spends more time on build and test than on inference, 750 tok/s gives you nothing back.
The bottom line
Ultrafast is honest and uncomfortable at the same time. Honest because it doesn't pretend to be a new model: it's the same Sol delivered faster, with real hardware engineering behind it. Uncomfortable because it showed up with vendor benchmarks, no price and invite-only access — on the same day Google put Gemini 3.7 Flash in GA in 160+ countries with a price on the table. One shipped a product, the other shipped a preview.
What changes for you today: nothing on the bill, a lot in your head. Speed became a SKU. Picking a model now has three axes — which model, how much it thinks, how fast it talks — each with its own cost. Start with the one that already has a price sheet: Sol in a real business, before dreaming about the tier nobody knows how to price.
{AI Engineer} — apaixonado por Laravel, arquitetura de software e construir produtos com impacto. Compartilho aqui tutoriais, descobertas e reflexões sobre o dia a dia de engenharia.
There is no shortage of content. What is missing is someone to untangle it: what matters now is how to implement it the right way. In the Clã you get that live, every week, with people who have already filtered out the noise.
Join the Clã