GPT-6 Sol and Luna, Opus 5.5, and Gemini 4: What Actually Changes in This Wave of Models
OpenAI's GPT-6 Sol and GPT-6 Luna and Anthropic's Claude Opus 5.5 shipped on the same day: September 22, ninety minutes apart. Two days later, the head of Google DeepMind got on stage to say Gemini 4 is already in post-training.
Three labs, one week. The whole timeline screaming "new model."
But if you look at the numbers calmly, this round has a trait the previous ones didn't: almost nothing here is about the model getting smarter. It's about the same intelligence getting cheaper. And that moves the math for anyone running agents in production far more than one extra benchmark point.
In this post you'll see what each release actually brings that's different, where the catch is in each one, what's already known about Gemini 4, and which model makes sense for each kind of workload.
TL;DR
- GPT-6 Sol: same intelligence as GPT-5.6 Sol at half the price ($2/$10 per million tokens), hallucinating a lot less.
- GPT-6 Luna: the volume model. $0.10/$0.50, but it slips a bit on agentic coding and burns more tokens per task.
- Claude Opus 5.5: delivers at Fable 5.1 level for $4/$20, 30% faster than Opus 5, with cache reads 60% cheaper. Ships with 4 breaking changes in the API.
- Gemini 4: not out yet. Pre-training confirmed in July, post-training now, prediction market betting on October.
The pattern of this round: price drops, intelligence stays
Start with the most honest data point there is on GPT-6 Sol. It's not OpenAI's. It's from Artificial Analysis, which runs the same tests on everyone.
Their verdict: GPT-6 Sol's Intelligence Index landed at the same level as GPT-5.6 Sol. Luna's, at the same level as GPT-5.6 Luna. With gains on some evals and regressions on others.
What changed was the cost per task:
| Model | Cost per task (Intelligence Index, max effort) |
|---|---|
| GPT-5.6 Sol | $1.99 |
| GPT-6 Sol | $1.06 |
| GPT-5.6 Luna | $0.18 |
| GPT-6 Luna | $0.07 |
On Anthropic's side, same story in different packaging. Opus 5.5 "performs at the level of Claude Fable 5.1 on most work" and costs 40% less than Opus 5 on typical workloads.
In other words: neither lab sold you a capability jump this week. Both sold you last month's frontier capability at half the price.
Sounds less exciting. It isn't.
Agents in production don't die from lack of intelligence. They die from cost. An agent loop that makes 40 calls per task and used to cost R$ 2 per run can now cost R$ 1. That's the difference between a feature that pencils out and one that stays a prototype.
And this is the pace: every week some lab touches the price table, a parameter default, or an API behavior. Keeping up with that on your own, with production code in your hands, is what wears you down the most. That's why at the Clã Beer and Code we take a release like this one and test it together, every week, live, on a real project.
GPT-6 Sol and Luna: the change that matters is hallucination
OpenAI positions the two like this: Sol is the everyday model, with more reasoning, and Luna is the cheap one, optimized for fast responses and high volume. According to TechCrunch, Luna targets "high-volume tasks with a clear objective, like summarizing documents, extracting information, or answering quick questions."
The prices, before and after (MacRumors):
| Model | Input (1M tokens) | Output (1M tokens) |
|---|---|---|
| GPT-5.6 Sol | $4 | $20 |
| GPT-6 Sol | $2 | $10 |
| GPT-5.6 Luna | $0.20 | $1.20 |
| GPT-6 Luna | $0.10 | $0.50 |
OpenAI attributes the cut to "improvements in caching and inference." It's not a promotion. It's infrastructure efficiency.
Now, my favorite data point of this round: hallucination.
On AA-Omniscience, a test that measures how often the model makes up an answer instead of admitting it doesn't know, Sol's hallucination rate dropped from 92% to 60%. Luna's, from 93% to 77% (Artificial Analysis). OpenAI itself says Sol makes "about half the errors" of GPT-5.6 Sol, reaching Astra's reliability.
60% is still high. But the drop is huge, and it lands squarely on RAG, extraction, and any flow where the model answers about data it doesn't have.
There's also a style change inherited from Astra: slightly shorter responses, less jargon and, according to OpenAI, "fewer misleading claims about its own coding work." Anyone who's caught an agent saying "all tests passed" without having run anything knows what that's worth.
The catch with Luna
Luna is cheap per token. But look at what Artificial Analysis measured:
- On the Coding Agent Index, Luna dropped 2 points (41). Sol went up 2 (57).
- Luna spent 51,000 output tokens on the index tasks, versus 41,000 for GPT-5.6 Luna.
- Both models regressed on GDPval-AA v2.1 at max effort, because they deliver shorter documents that "more often omit required elements."
Translation: if your Luna flow is agentic coding or document generation with a content checklist, don't swap blind. Run your eval first. For summarization and extraction, the swap is close to a no-brainer.
Claude Opus 5.5: Fable level at Opus pricing, with four breaking changes
Opus 5.5 is the first model in the 5.5 family. Model id: claude-opus-5-5.
The numbers against Opus 5, all from the official announcement:
| Benchmark | Opus 5 | Opus 5.5 |
|---|---|---|
| Terminal-Bench 4.0 | 52.3% | 66.4% |
| CursorBench 4.0 | 46.6% | 57.8% |
| FrontierCode v1.1 | 48.0% | 54.4% |
| OSWorld 2.0 | 74.0% | 81.8% |
| Humanity's Last Exam | 63.6% | 67.7% |
| GDPval-AA v2.1 (Elo) | 1708 | 1846 |
Fourteen points on Terminal-Bench in a point-five release is no small thing. And Artificial Analysis put Opus 5.5 at the top of the Intelligence Index, with 58 points.
But what I want you to look at is the price:
| Opus 5 | Opus 5.5 | |
|---|---|---|
| Input | $5 | $4 |
| Output | $25 | $20 |
| Cache write | $6.25 | $5 |
| Cache read | $0.50 | $0.20 |
Cache read dropped 60%. In an agentic session, where the entire context gets re-read every turn, that's the line that drives the bill, not input. We already ran that math in the Fable 5.1 vs Opus 5 comparison: with a warm cache, the cost ratio between models changes completely. With Opus 5.5 it tilts hard in its favor.
Anthropic also compares it to GPT-6 Astra: same level on FrontierCode at "about 20% of the cost per task," and on Terminal-Bench 4.0 at "about 40%." That's the lab's own number, so read it with the usual discount.
On top of that: it generates output more than 30% faster than Opus 5, and in the behavioral audit it tried to work around limits "about 85% less often" than Opus 5.
What breaks if you only swap the model id
Here's where the problem lives. The migration guide lists four breaking changes:
- Thinking can't be turned off.
thinking: {"type": "disabled"}andthinking: {"type": "enabled", "budget_tokens": N}return a 400 error. Adaptive thinking runs on every request. The only control now iseffort. - Forced tool use was removed.
tool_choicewith{"type": "any"}or{"type": "tool", ...}returns 400. Useautowith strict tool use or structured outputs. - Thinking blocks are locked to the conversation. The API only accepts a resent thinking block if everything that came before it stays identical. Edited an earlier turn, removed an injected reminder, rebuilt the system prompt, or did client-side compaction? Error. For accounts created on or after August 31, 2026, this is the default behavior.
- Computer use changed. On the API and on Google Cloud, the old
computer_20251124is rejected. The new one iscomputer_toolset_20260801.
And there's a silent change that throws no error at all, it just changes the result: the default effort is now medium. On Opus 5 it was high. If you migrate without setting it, you'll run with less reasoning and think the model got worse.
Before, on Opus 5, turning thinking off to save money:
response = client.messages.create(
model="claude-opus-5",
max_tokens=4096,
thinking={"type": "disabled"},
messages=[{"role": "user", "content": prompt}],
)
After, on Opus 5.5:
response = client.messages.create(
model="claude-opus-5-5",
max_tokens=16000, # agora cobre thinking + texto
output_config={"effort": "low"}, # thinking sempre ligado; effort é o controle
messages=[{"role": "user", "content": prompt}],
)
# a resposta começa com blocos de thinking: filtre por tipo
text = "".join(b.text for b in response.content if b.type == "text")
Two details that trip people up. First: max_tokens is now the ceiling for thinking plus text. If you were using 4096 without thinking, you may start getting truncated responses. The docs suggest starting at 64k for xhigh and max. Second: thinking is billed as output even when you don't see its text. A flow that ran without thinking will spend more output per request.
If you use Claude Code, you can let Claude itself do the migration with /claude-api migrate this project to claude-opus-5-5.
Sonnet 5.5 and Haiku 5.5 are coming "in the coming weeks," according to Anthropic.
Do not just follow the news — master it. Hands-on AI Engineering, live, every week, in the largest community in Brazil.
Join the ClãGemini 4: what's known so far and the forecast
Google is the only one of the three that didn't ship anything this week. And it knows it: the company's last flagship was the Gemini 3 series, in November 2025.
What's confirmed, in order:
- July 21, 2026: Google publicly confirmed it had started the "most ambitious pre-training to date" for Gemini 4, a rare confirmation made mid-training (MindStudio).
- July 23, 2026: on the earnings call, Sundar Pichai said Gemini 4 will use a "significantly larger" base model, with code and autonomous agents as priorities.
- This week: Koray Kavukcuoglu, in his first public appearance as head of Google DeepMind (The Information's AI Agenda Live), said Gemini 4 is early in post-training and that the goal is to ship "well before" the end of the year. In his words: "Our intention is to release an early version of the post-training as soon as possible, because we've already seen promising results" (AI Weekly).
And what the market is betting. On Polymarket, the odds of Gemini 4.0 shipping by each date, as of today:
| By | Probability |
|---|---|
| Sep 30, 2026 | 2% |
| Oct 31, 2026 | 75% |
| Nov 30, 2026 | 92% |
Important detail: this market resolves on Gemini 4.0 Flash being publicly available, not Pro.
My read. "An early version of the post-training as soon as possible" sounds like a preview, probably a Flash or a Pro with an experimental label, in October. Given the priorities Pichai stated, the focus will be code and agents. That's exactly where Opus 5.5 and GPT-6 Sol just raised the bar on price-performance. So the question that matters for Gemini 4 isn't "will it be smarter?" It's "will it be cheaper per solved task than Opus 5.5?" If it isn't, it shows up late.
This is a forecast, not a fact. Treat it as one.
Which one to use now
| Workload | Pick | Why |
|---|---|---|
| Coding agent, long-running task | Opus 5.5 | Best Terminal-Bench and CursorBench of the round, cheap cache read |
| Summarization, extraction, classification at volume | GPT-6 Luna | $0.10/$0.50, lower hallucination than its predecessor |
| General-purpose assistant, RAG | GPT-6 Sol | Half the price of 5.6, sharp drop in hallucination |
| Already on GPT-5.6 Sol | GPT-6 Sol | Same level, half the cost, low-risk swap |
| Already on Opus 5 | Opus 5.5, with a migration | Cheaper and better, but read the breaking changes first |
| Waiting for Gemini 4 | Don't wait | No official date; migrate what you can and reassess in October |
Limitations and things to watch
A lab benchmark is a lab benchmark. The Opus 5.5 numbers are Anthropic's. The line "outperforms Anthropic's models" is OpenAI's. The independent numbers I used are from Artificial Analysis, and even those are measured with their own harness. Terminal-Bench in Anthropic's announcement and Terminal-Bench on Artificial Analysis are not directly comparable.
Same intelligence doesn't mean same behavior. Sol and Luna got more concise. If your product depends on long, complete responses, the GDPval regression shows you can lose content. Your own eval, always.
Not everything is everywhere. On launch day, GPT-6 Sol and Luna were in ChatGPT Work, Codex, and the API, but not yet in Chat mode.
The Gemini forecast is a bet. Kavukcuoglu said "intention," not a date.
Quick FAQ
Is GPT-6 Sol better than GPT-6 Astra? No. Astra is still the top of OpenAI's lineup. Sol brings Astra's improvements to a lower price, with similar reliability on factuality, but it's not a replacement in raw capability. The full Astra comparison is in our comparison post.
Can I swap claude-opus-5 for claude-opus-5-5 and call it done?
Only if you don't turn thinking off, don't force tool use, don't edit the conversation history, and don't use computer use. Otherwise, you'll eat a 400. And set effort explicitly: the default dropped from high to medium.
Does Opus 5.5 replace Fable 5.1? For most work, according to Anthropic, yes, and at a much lower cost. For truly frontier tasks, Fable 5.1 is still the ceiling. Test it on your hardest cases before turning Fable off.
When does Gemini 4 come out? There's no official date. Google says "well before the end of the year," and Polymarket gives a 75% chance of a public Gemini 4.0 Flash by October 31.
Conclusion
This round didn't bring the model that changes everything. It brought something more useful: frontier capability got cheaper and more reliable. GPT-6 Sol hallucinates a lot less at half the price. Opus 5.5 delivers Fable level at Opus pricing, with a cache that makes long-running agents pencil out.
The next step in this race isn't who has the biggest number on the benchmark. It's who solves the task at the lowest total cost. Gemini 4 will be judged by that yardstick.
If you're still deciding between the top-tier models, it's worth reading the GPT-6 Astra vs Fable 5.1 comparison to see how the cache changes the math for an entire session.
{AI Engineer} — apaixonado por Laravel, arquitetura de software e construir produtos com impacto. Compartilho aqui tutoriais, descobertas e reflexões sobre o dia a dia de engenharia.
There is no shortage of content. What is missing is someone to untangle it: what matters now is how to implement it the right way. In the Clã you get that live, every week, with people who have already filtered out the noise.
Join the Clã