GPT-5.6 in Codex: Ultra mode buys 3 points, and Terra burns 2.6x more tokens to deliver worse results
Updated on August 15, 2026. The pricing table has been corrected: on July 30, 2026, OpenAI cut Luna by 80% and Terra by 20%. The numbers below are the current ones. The redone math is in GPT-5.6 Luna price dropped 80%.
Ten days ago, the #1 story on Hacker News wasn't a paper or a funding round. It was a sentence: GPT-5.6 Sol Ultra will be in Codex. 372 points, and the discussion caught fire somewhere between "this changes agentic coding" and "this is just a prompt alias."
Now it's a product. Since July 9, GPT-5.6 has been live in ChatGPT, the API and Codex, across all three tiers: Sol, Terra and Luna. And the launch brought two things that didn't exist during hype week: an official pricing table and the first independent tests. They tell a less pretty story than the marketing does, and the most useful part is right below.
TL;DR
- Status: launched on July 9, 2026, generally available in ChatGPT, the API and Codex. Ultra mode coordinates subagents by default on the top tier.
- Price per 1M tokens (July 30, 2026 pricing): Sol $5 input / $30 output; Terra $2 / $12; Luna $0.20 / $1.20. The map of when to use each tier is in the Sol, Terra and Luna comparison.
- OpenAI's number: 91.9% on Terminal-Bench 2.1 for Sol Ultra versus 88.8% for base Sol. Three points.
- The independent number that changes the decision: in CodeRabbit's tests, Sol passed 63.7% of the coding tasks; Terra passed only 40.7%, while burning 2.6x more tokens than Sol. Price per token is not price per task.
- What's technically different about Ultra: subagents trained to cooperate and exchange messages during the task, instead of running in parallel and in isolation like in Pro mode.
This post focuses on Ultra mode and the independent benchmarks of Sol Ultra inside Codex. If what you want is to decide between Sol, Terra and Luna in general (updated pricing, when to use each tier and the production math), the full guide is GPT-5.6 Sol, Terra or Luna: which one to use.
What Sol Ultra is (and why it sits "above" Sol)
Pay attention here, because the naming is confusing on purpose.
Sol is the strongest GPT-5.6 model, with two levers: max reasoning effort (it thinks deeper before answering) and ultra mode, which, per the official launch description, coordinates four subagents in parallel by default, capable of writing and running small programs to orchestrate tools without a manual script.
The "Sol Ultra" that runs in Codex is that ultra behavior packaged as a product experience. According to OpenAI, the difference from Pro mode (which runs agents in parallel, each on its own) is that here the subagents are trained to cooperate and communicate during the task. Instead of N independent attempts and a judge picking the best one, you get a team that splits the work and talks while it executes.
Thibault Sottiaux, who runs Codex at OpenAI, summed up the pitch with "save your hardest prompts somewhere." Translation: it's supposed to survive a long, hard task without falling apart halfway through.
What GPT-5.6 in Codex changes in practice
This is where it matters for people who write code, not people who read changelogs.
Codex is OpenAI's terminal and repository agent. With Sol Ultra inside Codex, subagent orchestration stopped being something you wire up by hand and became a button inside the tool. Three things change day to day:
- Long-horizon tasks without a babysitter. A refactor that touches ten files, a dependency migration, raising test coverage across an entire module: the kind of work where a solo agent loses the thread halfway. The team of subagents splits the work and holds context longer.
- Less glue on your side. Anyone who has built a subagent pipeline by hand knows half the work is writing the orchestrator. With it built in, you get time back to review the result instead of stitching the plumbing together.
- Terminal automation that can't fail. This is exactly the axis Terminal-Bench measures, and it's where the 91.9% number is trying to sell you confidence.
The honest caveat hasn't changed: none of this is model magic. It's orchestration. And orchestration is something you could already do with base Sol and a bit of engineering. Ultra packages it; it doesn't invent it. Knowing how to build that harness by hand is what separates people who flip a switch from people who design the system, and it's the kind of thing you learn by building alongside people who've already taken the hits: that's what we do every week, live, in the Clã Beer and Code.
The first independent tests (and the Terra trap)
This is the part that changes decisions, and it goes against most people's reflex.
CodeRabbit ran Sol and Terra on its own suite of coding and review tasks. The result:
- Coding tasks: Sol passed 63.7%, spending an average of 20,968 output tokens per task. Terra passed 40.7%, spending 55,594 tokens per task.
- Code review: Sol got 69.7% of actionable findings right (+7.4 points over their baseline). Terra landed below the baseline (-8.6 points).
Read that first bullet again, because it's counterintuitive: Terra costs less per token, but burns 2.6x more tokens to deliver a worse result. Price per token is not price per task. The "budget" tier can end up more expensive and wrong in agentic coding. It exists for a different workload profile (high volume, short tasks), not to replace Sol on the heavy lifting.
It's one test, from one vendor, on its own suite. It's not a verdict. But it's exactly the kind of independent verification that OpenAI's 91.9% still doesn't have.
Do not just follow the news — master it. Hands-on AI Engineering, live, every week, in the largest community in Brazil.
Join the ClãSol Ultra vs "regular" Sol: is it worth the bill?
This is the question that separates devs who think about architecture from devs who just want "the top one."
GPT-5.6 Sol costs $5 input and $30 output per 1M tokens. Ultra mode doesn't change the price list; it changes consumption. Several subagents cooperating means more calls, more context moving around, more tokens burned per task. The stated gain on Terminal-Bench is ~3 percentage points (88.8% to 91.9%) for a cost that can be several times higher.
Two cushions on the bill, confirmed at launch: GPT-5.6 shipped with explicit cache breakpoints and a 30-minute minimum lifetime on the prompt cache, which makes precisely the pattern of subagents re-reading the same context cheaper. It helps, but it doesn't flip the logic.
Do the engineer's math. Three points of accuracy on a critical refactor where a mistake costs a production incident? Worth it. Three points on a routine task base Sol already handles? You're paying flagship prices for mid-size work. A comment on Hacker News captured the corporate irony: management first praised whoever burned the most tokens, and weeks later asked everyone to cut costs. Ultra mode is beautiful until the invoice shows up. Routing by effort, with Ultra only where being right is worth more than the bill, is still the grown-up way to use this.
Sol Ultra vs Claude Code: deterministic subagents or not
Since Codex goes head to head with Claude Code in agentic coding, the technical comparison is worth making, and it's more interesting than "who has the bigger benchmark." If you're deciding between the two tools, the full comparison is in Claude Code or Codex.
The difference that came up in the HN discussion: Claude Code, in its heaviest mode, generates code to orchestrate the subagents deterministically. It builds a workflow you can inspect and reproduce. OpenAI's Ultra mode lets the model spawn subagents non-deterministically, deciding on the fly.
That's not an academic detail. It's the difference between an agent whose plan you can audit and an agent that picks its own path on every run. For anyone putting this in production, a deterministic workflow is easier to test, version and trust. Non-deterministic is more flexible, but you depend more on the luck of the model choosing well. Neither one is "the right one." They're different trade-offs, and you should know which one you're buying.
Healthy skepticism: is "ultra" a new model or a new prompt?
The bucket of cold water that saved the hype thread still stands after the launch.
The top-voted comment on Hacker News argued that "ultra" may not be a new architecture at all: it would be an alias for maximum effort plus an instruction for the model to use subagents proactively. The launch documentation describes the behavior (four coordinated agents by default), but OpenAI still hasn't published a detailed evaluation methodology for the mode, so the doubt remains legitimate.
If that's what it is, Sol Ultra is less a "top-tier model" and more a "mode of operation." Which isn't bad. It just changes how you read the marketing: you're not buying a better brain, you're buying an orchestration behavior that's on by default. And there's a good implication for you: if it's prompting + effort + subagents, you can reproduce a good chunk of the gain with base Sol and your own engineering. Ultra sells convenience, not a secret you couldn't build yourself.
What to do now
Access is open, so the right move stopped being "wait" and became measure. Before you swap your daily driver:
- Run Sol (without Ultra) on your real repo on a task your current agent already does. Compare result and cost per task, not benchmarks.
- Turn Ultra on only for the task that justifies it: a long refactor, a migration, the technical debt of an entire module. And look at the invoice afterward.
- Ignore the "Terra because it's cheaper" reflex for agentic coding. CodeRabbit's test shows why. If the goal is to cut cost at high volume, the right tier is Luna, and even then only after an eval.
Sol Ultra in Codex is good news for people who code. But it's a tool, not a miracle. Whoever treats agent orchestration as an architecture decision, and not as a "turbo mode" button, will get a lot more out of it than whoever just turns on the most expensive tier and prays.
FAQ
What is GPT-5.6 Sol Ultra?
It's the top tier of GPT-5.6 in ultra mode: Sol (OpenAI's strongest model) coordinating subagents trained to cooperate and exchange messages during the task, packaged to run inside Codex.
Can I already use GPT-5.6 in Codex? Yes. Since July 9, 2026, GPT-5.6 has been generally available in ChatGPT, the API and Codex, with Ultra mode on the top tier.
How much does GPT-5.6 cost? Per 1M tokens, on the pricing in effect since July 30, 2026: Sol $5 input / $30 output; Terra $2 / $12; Luna $0.20 / $1.20. Ultra mode doesn't change the unit price, but it increases consumption because it runs several subagents per task.
What's the difference between Sol Ultra and regular Sol?
Base Sol already reasons deeply with max effort. Ultra adds the orchestration of cooperating subagents. Stated gain from 88.8% to 91.9% on Terminal-Bench 2.1, at a much higher token cost.
Is it worth using Terra to save money in Codex? Based on CodeRabbit's data, no. Terra passed 40.7% of the tasks versus Sol's 63.7% and burned 2.6x more tokens in the process. In agentic coding, the middle tier ends up expensive and gets more wrong.
Is Sol Ultra better than Claude Code? It depends on what you value. Claude Code orchestrates subagents deterministically (an auditable workflow); Sol Ultra lets the model decide on the fly (non-deterministic). One is easier to test, the other more flexible.
{AI Engineer} — apaixonado por Laravel, arquitetura de software e construir produtos com impacto. Compartilho aqui tutoriais, descobertas e reflexões sobre o dia a dia de engenharia.
There is no shortage of content. What is missing is someone to untangle it: what matters now is how to implement it the right way. In the Clã you get that live, every week, with people who have already filtered out the noise.
Join the Clã