~/beer-and-code
▪ next event Workshop: Jev na Prática para Devs · 14 Oct · 19h — duração de 2h a 3h, ao vivo via Google Meet save your seat ›
~ / news / gpt-6-sol-luna-opus-5-5-gemini-4 $
News

GPT-6 Sol and Luna, Opus 5.5, and Gemini 4: What Actually Changes in This Wave of Models

LS Lucas Souza · · 13 min read
GPT-6 Sol and Luna, Opus 5.5, and Gemini 4: What Actually Changes in This Wave of Models

OpenAI's GPT-6 Sol and GPT-6 Luna and Anthropic's Claude Opus 5.5 shipped on the same day: September 22, ninety minutes apart. Two days later, the head of Google DeepMind got on stage to say Gemini 4 is already in post-training.

Three labs, one week. The whole timeline screaming "new model."

But if you look at the numbers calmly, this round has a trait the previous ones didn't: almost nothing here is about the model getting smarter. It's about the same intelligence getting cheaper. And that moves the math for anyone running agents in production far more than one extra benchmark point.

In this post you'll see what each release actually brings that's different, where the catch is in each one, what's already known about Gemini 4, and which model makes sense for each kind of workload.

TL;DR

  • GPT-6 Sol: same intelligence as GPT-5.6 Sol at half the price ($2/$10 per million tokens), hallucinating a lot less.
  • GPT-6 Luna: the volume model. $0.10/$0.50, but it slips a bit on agentic coding and burns more tokens per task.
  • Claude Opus 5.5: delivers at Fable 5.1 level for $4/$20, 30% faster than Opus 5, with cache reads 60% cheaper. Ships with 4 breaking changes in the API.
  • Gemini 4: not out yet. Pre-training confirmed in July, post-training now, prediction market betting on October.

The pattern of this round: price drops, intelligence stays

Start with the most honest data point there is on GPT-6 Sol. It's not OpenAI's. It's from Artificial Analysis, which runs the same tests on everyone.

Their verdict: GPT-6 Sol's Intelligence Index landed at the same level as GPT-5.6 Sol. Luna's, at the same level as GPT-5.6 Luna. With gains on some evals and regressions on others.

What changed was the cost per task:

Model Cost per task (Intelligence Index, max effort)
GPT-5.6 Sol $1.99
GPT-6 Sol $1.06
GPT-5.6 Luna $0.18
GPT-6 Luna $0.07

On Anthropic's side, same story in different packaging. Opus 5.5 "performs at the level of Claude Fable 5.1 on most work" and costs 40% less than Opus 5 on typical workloads.

In other words: neither lab sold you a capability jump this week. Both sold you last month's frontier capability at half the price.

Sounds less exciting. It isn't.

Agents in production don't die from lack of intelligence. They die from cost. An agent loop that makes 40 calls per task and used to cost R$ 2 per run can now cost R$ 1. That's the difference between a feature that pencils out and one that stays a prototype.

And this is the pace: every week some lab touches the price table, a parameter default, or an API behavior. Keeping up with that on your own, with production code in your hands, is what wears you down the most. That's why at the Clã Beer and Code we take a release like this one and test it together, every week, live, on a real project.

GPT-6 Sol and Luna: the change that matters is hallucination

OpenAI positions the two like this: Sol is the everyday model, with more reasoning, and Luna is the cheap one, optimized for fast responses and high volume. According to TechCrunch, Luna targets "high-volume tasks with a clear objective, like summarizing documents, extracting information, or answering quick questions."

The prices, before and after (MacRumors):

Model Input (1M tokens) Output (1M tokens)
GPT-5.6 Sol $4 $20
GPT-6 Sol $2 $10
GPT-5.6 Luna $0.20 $1.20
GPT-6 Luna $0.10 $0.50

OpenAI attributes the cut to "improvements in caching and inference." It's not a promotion. It's infrastructure efficiency.

Now, my favorite data point of this round: hallucination.

On AA-Omniscience, a test that measures how often the model makes up an answer instead of admitting it doesn't know, Sol's hallucination rate dropped from 92% to 60%. Luna's, from 93% to 77% (Artificial Analysis). OpenAI itself says Sol makes "about half the errors" of GPT-5.6 Sol, reaching Astra's reliability.

60% is still high. But the drop is huge, and it lands squarely on RAG, extraction, and any flow where the model answers about data it doesn't have.

There's also a style change inherited from Astra: slightly shorter responses, less jargon and, according to OpenAI, "fewer misleading claims about its own coding work." Anyone who's caught an agent saying "all tests passed" without having run anything knows what that's worth.

The catch with Luna

Luna is cheap per token. But look at what Artificial Analysis measured:

  • On the Coding Agent Index, Luna dropped 2 points (41). Sol went up 2 (57).
  • Luna spent 51,000 output tokens on the index tasks, versus 41,000 for GPT-5.6 Luna.
  • Both models regressed on GDPval-AA v2.1 at max effort, because they deliver shorter documents that "more often omit required elements."

Translation: if your Luna flow is agentic coding or document generation with a content checklist, don't swap blind. Run your eval first. For summarization and extraction, the swap is close to a no-brainer.

Claude Opus 5.5: Fable level at Opus pricing, with four breaking changes

Opus 5.5 is the first model in the 5.5 family. Model id: claude-opus-5-5.

The numbers against Opus 5, all from the official announcement:

Benchmark Opus 5 Opus 5.5
Terminal-Bench 4.0 52.3% 66.4%
CursorBench 4.0 46.6% 57.8%
FrontierCode v1.1 48.0% 54.4%
OSWorld 2.0 74.0% 81.8%
Humanity's Last Exam 63.6% 67.7%
GDPval-AA v2.1 (Elo) 1708 1846

Fourteen points on Terminal-Bench in a point-five release is no small thing. And Artificial Analysis put Opus 5.5 at the top of the Intelligence Index, with 58 points.

But what I want you to look at is the price:

Opus 5 Opus 5.5
Input $5 $4
Output $25 $20
Cache write $6.25 $5
Cache read $0.50 $0.20

Cache read dropped 60%. In an agentic session, where the entire context gets re-read every turn, that's the line that drives the bill, not input. We already ran that math in the Fable 5.1 vs Opus 5 comparison: with a warm cache, the cost ratio between models changes completely. With Opus 5.5 it tilts hard in its favor.

Anthropic also compares it to GPT-6 Astra: same level on FrontierCode at "about 20% of the cost per task," and on Terminal-Bench 4.0 at "about 40%." That's the lab's own number, so read it with the usual discount.

On top of that: it generates output more than 30% faster than Opus 5, and in the behavioral audit it tried to work around limits "about 85% less often" than Opus 5.

What breaks if you only swap the model id

Here's where the problem lives. The migration guide lists four breaking changes:

  1. Thinking can't be turned off. thinking: {"type": "disabled"} and thinking: {"type": "enabled", "budget_tokens": N} return a 400 error. Adaptive thinking runs on every request. The only control now is effort.
  2. Forced tool use was removed. tool_choice with {"type": "any"} or {"type": "tool", ...} returns 400. Use auto with strict tool use or structured outputs.
  3. Thinking blocks are locked to the conversation. The API only accepts a resent thinking block if everything that came before it stays identical. Edited an earlier turn, removed an injected reminder, rebuilt the system prompt, or did client-side compaction? Error. For accounts created on or after August 31, 2026, this is the default behavior.
  4. Computer use changed. On the API and on Google Cloud, the old computer_20251124 is rejected. The new one is computer_toolset_20260801.

And there's a silent change that throws no error at all, it just changes the result: the default effort is now medium. On Opus 5 it was high. If you migrate without setting it, you'll run with less reasoning and think the model got worse.

Before, on Opus 5, turning thinking off to save money:

response = client.messages.create(
    model="claude-opus-5",
    max_tokens=4096,
    thinking={"type": "disabled"},
    messages=[{"role": "user", "content": prompt}],
)

After, on Opus 5.5:

response = client.messages.create(
    model="claude-opus-5-5",
    max_tokens=16000,  # agora cobre thinking + texto
    output_config={"effort": "low"},  # thinking sempre ligado; effort é o controle
    messages=[{"role": "user", "content": prompt}],
)

# a resposta começa com blocos de thinking: filtre por tipo
text = "".join(b.text for b in response.content if b.type == "text")

Two details that trip people up. First: max_tokens is now the ceiling for thinking plus text. If you were using 4096 without thinking, you may start getting truncated responses. The docs suggest starting at 64k for xhigh and max. Second: thinking is billed as output even when you don't see its text. A flow that ran without thinking will spend more output per request.

If you use Claude Code, you can let Claude itself do the migration with /claude-api migrate this project to claude-opus-5-5.

Sonnet 5.5 and Haiku 5.5 are coming "in the coming weeks," according to Anthropic.

▪ Clã Beer and Code

Do not just follow the news — master it. Hands-on AI Engineering, live, every week, in the largest community in Brazil.

Join the Clã

Gemini 4: what's known so far and the forecast

Google is the only one of the three that didn't ship anything this week. And it knows it: the company's last flagship was the Gemini 3 series, in November 2025.

What's confirmed, in order:

  • July 21, 2026: Google publicly confirmed it had started the "most ambitious pre-training to date" for Gemini 4, a rare confirmation made mid-training (MindStudio).
  • July 23, 2026: on the earnings call, Sundar Pichai said Gemini 4 will use a "significantly larger" base model, with code and autonomous agents as priorities.
  • This week: Koray Kavukcuoglu, in his first public appearance as head of Google DeepMind (The Information's AI Agenda Live), said Gemini 4 is early in post-training and that the goal is to ship "well before" the end of the year. In his words: "Our intention is to release an early version of the post-training as soon as possible, because we've already seen promising results" (AI Weekly).

And what the market is betting. On Polymarket, the odds of Gemini 4.0 shipping by each date, as of today:

By Probability
Sep 30, 2026 2%
Oct 31, 2026 75%
Nov 30, 2026 92%

Important detail: this market resolves on Gemini 4.0 Flash being publicly available, not Pro.

My read. "An early version of the post-training as soon as possible" sounds like a preview, probably a Flash or a Pro with an experimental label, in October. Given the priorities Pichai stated, the focus will be code and agents. That's exactly where Opus 5.5 and GPT-6 Sol just raised the bar on price-performance. So the question that matters for Gemini 4 isn't "will it be smarter?" It's "will it be cheaper per solved task than Opus 5.5?" If it isn't, it shows up late.

This is a forecast, not a fact. Treat it as one.

Which one to use now

Workload Pick Why
Coding agent, long-running task Opus 5.5 Best Terminal-Bench and CursorBench of the round, cheap cache read
Summarization, extraction, classification at volume GPT-6 Luna $0.10/$0.50, lower hallucination than its predecessor
General-purpose assistant, RAG GPT-6 Sol Half the price of 5.6, sharp drop in hallucination
Already on GPT-5.6 Sol GPT-6 Sol Same level, half the cost, low-risk swap
Already on Opus 5 Opus 5.5, with a migration Cheaper and better, but read the breaking changes first
Waiting for Gemini 4 Don't wait No official date; migrate what you can and reassess in October

Limitations and things to watch

A lab benchmark is a lab benchmark. The Opus 5.5 numbers are Anthropic's. The line "outperforms Anthropic's models" is OpenAI's. The independent numbers I used are from Artificial Analysis, and even those are measured with their own harness. Terminal-Bench in Anthropic's announcement and Terminal-Bench on Artificial Analysis are not directly comparable.

Same intelligence doesn't mean same behavior. Sol and Luna got more concise. If your product depends on long, complete responses, the GDPval regression shows you can lose content. Your own eval, always.

Not everything is everywhere. On launch day, GPT-6 Sol and Luna were in ChatGPT Work, Codex, and the API, but not yet in Chat mode.

The Gemini forecast is a bet. Kavukcuoglu said "intention," not a date.

Quick FAQ

Is GPT-6 Sol better than GPT-6 Astra? No. Astra is still the top of OpenAI's lineup. Sol brings Astra's improvements to a lower price, with similar reliability on factuality, but it's not a replacement in raw capability. The full Astra comparison is in our comparison post.

Can I swap claude-opus-5 for claude-opus-5-5 and call it done? Only if you don't turn thinking off, don't force tool use, don't edit the conversation history, and don't use computer use. Otherwise, you'll eat a 400. And set effort explicitly: the default dropped from high to medium.

Does Opus 5.5 replace Fable 5.1? For most work, according to Anthropic, yes, and at a much lower cost. For truly frontier tasks, Fable 5.1 is still the ceiling. Test it on your hardest cases before turning Fable off.

When does Gemini 4 come out? There's no official date. Google says "well before the end of the year," and Polymarket gives a 75% chance of a public Gemini 4.0 Flash by October 31.

Conclusion

This round didn't bring the model that changes everything. It brought something more useful: frontier capability got cheaper and more reliable. GPT-6 Sol hallucinates a lot less at half the price. Opus 5.5 delivers Fable level at Opus pricing, with a cache that makes long-running agents pencil out.

The next step in this race isn't who has the biggest number on the benchmark. It's who solves the task at the lowest total cost. Gemini 4 will be judged by that yardstick.

If you're still deciding between the top-tier models, it's worth reading the GPT-6 Astra vs Fable 5.1 comparison to see how the cache changes the math for an entire session.

Lucas Souza
Written by
Lucas Souza

{AI Engineer} — apaixonado por Laravel, arquitetura de software e construir produtos com impacto. Compartilho aqui tutoriais, descobertas e reflexões sobre o dia a dia de engenharia.

▪ Clã Beer and Code

There is no shortage of content. What is missing is someone to untangle it: what matters now is how to implement it the right way. In the Clã you get that live, every week, with people who have already filtered out the noise.

Join the Clã
Meet the Clã Beer and Code
playing