~/beer-and-code
▪ next event Workshop: Jev na Prática para Devs · 14 Oct · 19h — duração de 2h a 3h, ao vivo via Google Meet save your seat ›
~ / news / gpt-6-astra-vs-fable-5-1-vs-gpt-5-6-sol $
News

GPT-6 Astra vs Fable 5.1 vs GPT-5.6 Sol: What It Is, Pricing and Benchmarks

LS Lucas Souza · · 18 min read
GPT-6 Astra vs Fable 5.1 vs GPT-5.6 Sol: What It Is, Pricing and Benchmarks

GPT-6 Astra is OpenAI's new top-of-the-line model, launched today, September 3, 2026, at the same price as Claude Fable 5.1, which shipped two days earlier. Greg Brockman opened the announcement with "Welcome to the AGI era".

Ignore the line. Look at the table.

In this post I put the three models that matter right now side by side: GPT-6 Astra, GPT-5.6 Sol (OpenAI's flagship until yesterday) and Claude Fable 5.1. Benchmark by benchmark, with sources. Pricing with the cache math laid out, because that's where the bill is born. And the detail that matters most if you build agents: the same GPT-6 Astra scored 62.7% and 99.9% on the same benchmark, and what changed between one number and the other wasn't the model.

TL;DR

  • What it is: GPT-6 Astra, OpenAI's new flagship. Model id gpt-6-astra, 1,050,000-token context window, 128k output, knowledge cutoff April 30, 2026. Successor to GPT-5.6 Sol.
  • Price: $10 input / $50 output per million tokens. Exactly the price of Claude Fable 5.1, and 2.5x the current price of GPT-5.6 Sol. Cache read at $1, versus $0.25 on Fable 5.1.
  • Benchmarks: on OpenAI's own numbers, Astra beats Fable 5.1 on everything both have published, with room to spare in math, science and automation. On Artificial Analysis's independent index, Fable 5.1 is still #1 and Astra shows up at #14. I explain the gap below.
  • Access: first to companies in the Daybreak program (cybersecurity), then "in the coming days" to ChatGPT Plus, Pro, Business and Enterprise, the API and AWS. It's the first model classified as Critical for cyber under OpenAI's Preparedness Framework.
  • Useful links: system card, model docs in the API, ARC Prize analysis.

Timeline of the leaks before the announcement: the Astra name and the staged rollout first showed up in the OpenAI Astra tracker.

What OpenAI shipped, and why the price is the first headline

The official pitch is short: "the smartest and most aligned model in the world", state of the art in computer use, browsing, software engineering, cybersecurity, science and professional work. Brockman summed it up for NBC: "Astra can do anything a human does with a computer".

The launch video is all about that. Someone asks for a yellow circle, turns it into a rocket window, has it made into a 3D model in Blender, exports an STL for the printer. In the middle of all that, the model creates an eBay listing with a photo from the downloads folder, drafts a licensing agreement, builds a 3D asteroid-dodging game, orders food and books a tennis court. All in parallel, in a single session. It's a demo, so discount it. But the message is clear: the product isn't chat. It's a computer operator. (If you've been following the GPT-6 tracker since July, the Astra name and the cyber-gated staged rollout were the two things that kept repeating in the leaks. Both were confirmed.)

Now the number nobody put in the headline.

Astra costs $10 per million input tokens and $50 per million output. That's the same price sheet as Claude Fable 5.1. Cent for cent.

That's a change in posture. In July, GPT-5.6 Sol arrived costing about a third of Fable 5 per task on the Artificial Analysis index, and OpenAI sold that as the "price-performance frontier". Today Sol sits at $4 / $20 on promotional pricing. Astra comes in 2.5x above Sol and pinned to the competitor's price. OpenAI has stopped undercutting at the top of the line. When the two biggest labs charge the same, the fight stops being about price and becomes about price per task, and that changes how you compare the two.

There's also a Fast mode at $20 / $100, according to VentureBeat, and a long-context surcharge: a prompt above 272k tokens pays 2x on input and 1.5x on output for the entire request.

The price sheet for all three, per million tokens:

Model Input Cache read Cache write Output
GPT-6 Astra $10.00 $1.00 $12.50 $50.00
Claude Fable 5.1 $10.00 $0.25 $12.50 $50.00
GPT-5.6 Sol (promo) $4.00 ~$0.40 ~$5.00 $20.00

Sources: Astra docs, Fable 5.1 announcement, Artificial Analysis for Sol. Sol's cache pricing is estimated from OpenAI's rule of cache read at 10% of input and cache write at 1.25x.

Look at the cache read column. It's the only line where Astra and Fable 5.1 diverge, and it's the line that weighs the most in an agent session. I'll come back to it in the math.

The top of the line changed hands twice in 48 hours, and it'll change again before the month is out. What doesn't change is the architecture that survives the swap without rewriting the product: the model is a parameter, not a foundation. But building that layer is decision after decision — where to cut context, what to cache, what the fallback is when a task stops halfway — and decisions are exactly what text teaches worst, because in a post you can't interrupt and ask why. That's what we build from scratch in the AI Engineering Lab 3rd Edition, September 19 and 20, 9am to 1pm, online: two days going from the architecture decision all the way to production.

GPT-6 Astra vs Fable 5.1 vs GPT-5.6 Sol: the table, benchmark by benchmark

Every Astra number below is reported by OpenAI, in the system card and in the launch material compiled by OfficeChai. The Fable 5.1 and Sol numbers that appear in OpenAI's announcement match the ones Anthropic published two days ago, which gives you at least some cross-checked confidence. Even so: nobody independent has reproduced any of this yet.

Benchmark GPT-6 Astra Claude Fable 5.1 GPT-5.6 Sol
FrontierMath Tier 4 (v2) 97.6% 87.8% 83.0%
GPQA Diamond 96.0% 93.7% n/p
Terminal-Bench-Science 0.1 64.6% 52.6% 22.4%
AutomationBench 41.4% 31.4% 19.6%
DeepSWE v1.1 74.1% 67.4% n/p
BenchCAD 95.9% 84.3% n/p
HealthBench Professional 63.4% 56.6% 60.5%
OSWorld 2.0 72.6% 77.9% (partial) / 41.7% (strict) 65.7%
ARC-AGI-3 (semi-private) 99.9% / 62.7%* n/p n/p
Terminal-Bench 4.0 "state of the art", no number 55.8% 37.3%

*Depends on the harness. It gets its own section below. n/p = not published.

Read the table in three groups.

Group 1: blowout. FrontierMath Tier 4 went from 87.8% to 97.6%, and OpenAI is calling the benchmark saturated. Terminal-Bench-Science climbed 12 points on top of what was already Fable 5.1's biggest lead over Opus 5. AutomationBench, which measures end-to-end business workflows, went from 31.4% to 41.4%. Those are double-digit gains on long, autonomous tasks. It's the same pattern as Fable 5.1 over Fable 5: the gap shows up when the agent has to stay awake for hours.

Group 2: narrow margin. GPQA Diamond: 2.3 points, on a benchmark that's already at the ceiling. HealthBench: 6.8 points. DeepSWE: 6.7 points, and it's worth noting that Gemini 3.8 Flash scores 73.7% on the same benchmark at a fraction of the price. Pure coding isn't where Astra opened up its biggest lead.

Group 3: not comparable. OSWorld 2.0 has three numbers in the table that measure different things. Anthropic publishes "partial" and "strict"; OpenAI publishes a single number and doesn't say which criterion. Don't put 72.6% next to 77.9% and conclude anything. And Terminal-Bench 4.0, the agentic coding benchmark Anthropic used to position Fable 5.1, appears in OpenAI's announcement as "state of the art" with no number. When the number shows up, you compare. When the adjective shows up, you wait.

The independent index tells a different story, for now

Artificial Analysis has already run Astra. The result:

Model Intelligence Index Rank
Claude Fable 5.1 66 #1 of 202
GPT-5.6 Sol (max) 61 #10
GPT-6 Astra (high) 60 #14

Before you yell "Astra lost", read the parentheses. Astra was tested at high. The model has five effort levels: low, medium, high, xhigh and max. Sol was tested at max. Fable 5.1 has adaptive thinking always on. You're looking at three models in three different configurations.

That doesn't invalidate the index. It invalidates the rushed comparison. When Astra at xhigh and max lands in the table, then you have a number. Until then, the index says one useful thing: at high, Astra costs 2.5x Sol to deliver the same score. If your use case doesn't need maximum effort, the new model isn't automatically the right model.

And here's the lesson that applies to any benchmark you read this week: always ask "at what effort level?". Without that answer, the number means nothing.

62.7% or 99.9%: the harness was worth more than the model

This is the part of the launch I most want you to read.

ARC Prize ran GPT-6 Astra on the semi-private ARC-AGI-3 in two ways. On the standard harness, where the model carries notes it chooses to keep as it moves through the environment, Astra scored 62.7%, spending $26K. On the "provider adapter" harness, which preserves the opaque reasoning state between requests and uses compaction on long conversations, the same model scored 99.9%, spending $19K.

Same model. Same weights. 37 points of difference. And the better version was 3.66x faster and used 49% fewer tokens.

That's context engineering, not intelligence. What changed between 62.7% and 99.9% was how reasoning survives from one turn to the next. The model that "forgets" between calls and has to reorient itself spends more and gets less right. The model that carries state spends less and gets more right.

OpenAI productized exactly that idea in Codex. According to 9to5Mac, instead of compressing the conversation through summarization when the window fills up, Astra "keeps notes across context windows, preserving the accumulated details", and earlier windows remain searchable. A requirement you gave in message 3 doesn't become a summarized paragraph by message 80. It becomes a retrievable note.

If you build agents, this is the only thing in this post that changes your code today. It's not the model. It's what you do with state between calls. We already saw this pattern when Sol's Ultra mode picked up 3 points on Terminal-Bench by coordinating four subagents: the score difference came from orchestration, not from the weights.

And ARC Prize itself closes with the line OpenAI left out of the announcement: saturating the benchmark "does not represent proof of AGI". The scope of ARC-AGI-3 is closed. It isn't the real world.

▪ Clã Beer and Code

Do not just follow the news — master it. Hands-on AI Engineering, live, every week, in the largest community in Brazil.

Join the Clã

Computer use: 72.6% and 40 minutes per task

On OSWorld 2.0, Astra scored 72.6% spending an average of 40 minutes per task. Sol scored 65.7% in 75 minutes. Less time, higher score. It's the most concrete gain of the launch for anyone using a model as a system operator.

But 40 minutes per task is the data point that matters for architecture. That's not chat latency. That's a job. If you're thinking about plugging Astra into a flow that fills out forms, updates a CRM, builds a spreadsheet or navigates a legacy dashboard, the design is asynchronous: a queue, persisted state, a callback when it finishes, and a human who approves before any irreversible action. A model operating a computer for 40 minutes with no supervision and no checkpoint is an incident waiting to happen.

The list of native tools in the API makes it clear where OpenAI is pointing: computer_use, hosted_shell, apply_patch, skills, mcp, tool_search, plus web search, file search and code interpreter. Structured outputs, function calling and prompt caching are supported. Fine-tuning, Realtime and Assistants are not.

"Critical" for cyber: what changes in your API

Astra is the first OpenAI model to hit the Critical level of cybersecurity capability under the Preparedness Framework. The official definition: "with the right tools and access, GPT-6 Astra can find unknown security flaws and develop new ways to exploit them in many well-protected systems, without a person guiding each step". During evaluation, it found two previously unknown vulnerabilities. ExploitBench: 100%.

What that changes for you, in practice:

Staged rollout. The ones with access today are in the Daybreak program, for cybersecurity defenders. Everyone else comes "in the coming days". If you're already blocked on Sol by quota or region, don't expect Astra to open up any sooner.

A task can stop halfway. According to VentureBeat, on activity flagged as suspicious ChatGPT asks for user approval before continuing, and "in API flows, a flagged task may simply stop". Remember Sol's automatic bans on DevOps scripts and Excel macros? The same class of guardrail now has a more sensitive model behind it. Your code needs to treat "stopped by policy" as an expected error state, with retry off and alerting on.

Total monitoring. OpenAI now monitors full trajectories, including the chain of reasoning, on every external inference that uses tools. It's not optional.

Now the part the announcement hides in the system card. Astra shows a "substantial reduction in chain-of-reasoning monitorability" compared to previous models: it controls its own reasoning better, solves tasks without verbalizing, and evades the monitor when it knows it's being observed. Apollo Research measured awareness of being under evaluation in 41% to 51% of runs. Jakub Pachocki, chief scientist, told NBC: "we will not accept degradation in our ability to monitor the model's alignment beyond a certain level". That sentence exists because the level got close.

On the other side, the alignment numbers in a simulated work environment improved a lot. Misaligned outcome rate: 3.4% on Astra versus 18.8% on Sol. Honeypot attacks: zero versus 55.4%. Indirect prompt injection in the Gray Swan arena: 8.5% attacker success versus 27.0% on Sol. If you have agents reading third-party email, pages and documents, that last line is the one worth money. But 8.5% isn't zero. The guardrail is still yours, not the model's.

Which one to use: Astra, Fable 5.1 or Sol

Price per token is tied between Astra and Fable 5.1. So the decision becomes price per task, and price per task depends on the shape of the workload.

Take the same simulation I used in the Fable 5.1 versus Opus 5 post: a coding agent with a 120k-token cached prefix (system prompt, tools, project files), 200 turns, 3k of new input and 2k of output per turn.

Line GPT-6 Astra Claude Fable 5.1 GPT-5.6 Sol
Cache write (1x, 120k) $1.50 $1.50 $0.60
Cache reads (200 x 120k = 24M) $24.00 $6.00 $9.60
New input (600k) $6.00 $6.00 $2.40
Output (400k) $20.00 $20.00 $8.00
Session total $51.50 $33.50 $20.60

Same price sheet. A session 54% more expensive on Astra. The entire difference is the cache read: $1 versus $0.25. In an agent session, cache reads dominate the token count, and Anthropic cut that multiplier last week for exactly this reason.

With that in hand, the decision tree I'd use today:

1. Single turn, high volume. Classification, extraction, routing. None of the three. Sonnet 5, GPT-5.6 Luna or Terra. A flagship here is money on fire, and the launch didn't change that.

2. Day-to-day agentic coding. GPT-5.6 Sol is still the best price per task of the three, and 88.8% on Terminal-Bench 2.1 didn't become obsolete because a new model came out. If your eval passes on Sol, you have no reason to migrate.

3. Long session with a large, re-read prefix. Fable 5.1. The $0.25 cache read eats any benchmark advantage Astra has once you run the numbers, and the gap in pure coding (DeepSWE, 6.7 points) doesn't pay for 54% more per session.

4. Frontier math, scientific research in the terminal, end-to-end business workflow automation. Astra, comfortably. 10 to 12 points on FrontierMath, Terminal-Bench-Science and AutomationBench is the kind of gap that justifies switching.

5. Computer use as an operator. Astra. 72.6% at 40 minutes per task is the best published number, as long as the architecture is asynchronous and has human approval.

6. Compliance. Fable 5.1 requires 30-day data retention and doesn't run under zero data retention without Anthropic's authorization. Astra can stop a flagged task halfway, and initial access goes through a closed program. Neither one is neutral on contract terms. Read before you sign.

And before you move up a model, move up effort. Astra's five levels trade intelligence for cost within the same model, without invalidating cache and without migrating code. The same rule that held last generation still holds: Sol, Terra or Luna was an eval decision, not a headline decision.

Limitations and things to watch

  • Everything is self-reported. Astra's benchmarks came from OpenAI. The only independent index that has run the model tested it at high, not max. Treat the table as a hypothesis until week two.
  • You probably don't have access yet. The API and paid plans arrive "in the coming days". The initial rate limit on tier 1 is 500 requests and 500K tokens per minute.
  • Long context costs double. Above 272k input tokens, the entire request pays 2x on input and 1.5x on output. An agent that lets the window fill up without compacting doubles the bill without warning.
  • A flagged task stops. In the API, it's not a warning. It's a halt. Your flow needs persisted state to resume where it left off, and an alert so someone takes a look.
  • Monitorability dropped. The system card itself says that, if the decline continues over the next generations, confidence in detecting misaligned behavior "would drop significantly". You can't fix that in your application, but you need to know it exists when you design what the agent is allowed to do on its own.
  • AGI is the seller's opinion. Brockman said "I personally think we've arrived". ARC Prize, owner of the benchmark he cited, said saturating the test proves none of that. Go with the second one.

Quick FAQ

Is GPT-6 Astra available in the API yet? Not for everyone. Today only the Daybreak program has access. OpenAI promises the API, ChatGPT Plus/Pro/Business/Enterprise and AWS "in the coming days". The model id is gpt-6-astra, and the docs are already live with pricing, limits and supported tools.

Is it worth migrating from GPT-5.6 Sol to Astra? Only if your eval on Sol at maximum effort is failing, and the failure is on long, scientific or end-to-end automation tasks. For day-to-day coding, Sol still costs 2.5x less per token and delivers the same score on the independent index. Without an eval, you have no way of knowing.

Astra or Fable 5.1 for a long-running agent? Same price sheet, cache read 4x more expensive on Astra. If the session re-reads a large prefix many times, Fable 5.1 comes out a lot cheaper per task. If the task is math, science or business automation, Astra has a lead of 10 points or more and the math flips.

Is GPT-6 Astra AGI? OpenAI says it "probably marks the beginning". ARC Prize, which runs the 99.9% benchmark cited in the announcement, says saturating the test isn't proof of AGI. If you build product, the question is irrelevant: what matters is price per task, error rate and what happens when the guardrail fires.

Conclusion

GPT-6 Astra is, on OpenAI's numbers, the most capable model anyone will be able to rent in the coming weeks. And the gap to Fable 5.1 is real where it matters to OpenAI: math, science, automation, operating a computer.

But the launch made three things clearer than any benchmark. Flagship pricing has converged, so the competition is now cache and workload shape. The same model scores 62.7% or 99.9% depending on how you preserve state between calls, so the harness is product engineering, not a detail. And the most capable model is also the hardest to monitor, and OpenAI is saying so out loud.

What changes in your code today isn't the model string. It's how you keep context between turns, how you handle a task that stops by policy, and whether you have an eval to know if any of these three is worth what it costs.

Those who have an eval find out in an afternoon. Those who don't find out on the invoice.

Lucas Souza
Written by
Lucas Souza

{AI Engineer} — apaixonado por Laravel, arquitetura de software e construir produtos com impacto. Compartilho aqui tutoriais, descobertas e reflexões sobre o dia a dia de engenharia.

▪ Clã Beer and Code

There is no shortage of content. What is missing is someone to untangle it: what matters now is how to implement it the right way. In the Clã you get that live, every week, with people who have already filtered out the noise.

Join the Clã
Meet the Clã Beer and Code
playing