Grok 4.7 Answers in 0.85s and Fable 5.1 in 298s: Where Each One Pays Off
Grok 4.7 costs $2 per million input tokens. Claude Fable 5.1 costs $10.
Five times cheaper.
And yet, on cost per completed task, it loses to GPT-6 Astra — which also charges $10.
That's the kind of number that only shows up when you stop staring at the pricing table and start looking at the final bill. This post takes Grok 4.7, released by xAI on September 21, 2026, and answers the only question that matters if you build product: in which scenario is this model the right pick — and in which ones is it the expensive pick dressed up as the cheap one.
Spoiler: there's one scenario where it has no competition. And it isn't writing code.
TL;DR
- What it is: Grok 4.7, xAI's (now SpaceXAI LLC) frontier model and the successor to Grok 4.6. Text and image in, text out.
- Context and cutoff: 500k tokens, knowledge through May 2026.
- Cost/Access: $2 input and $6 output per million, cached input at $0.50. Above 200k prompt tokens, everything doubles. Available on the xAI API, in Cursor, in Grok Build, and on the model routers.
- Where it wins: latency. First token in 0.85 seconds, versus 298 seconds for Fable 5.1 and 259 for GPT-6 Astra at their maximum reasoning settings.
- Where it loses: long agentic tasks. On Terminal-Bench 4.0 it scores 38.0% by xAI's own measurement and 26% by independent measurement — against 60% for GPT-6 Astra.
- Useful link: official announcement | models and pricing docs
What xAI shipped on September 21
The announcement is blunt: bigger base model, longer reinforcement learning cycle, trained with more weight on problems that take many hours to solve. xAI says it got "better at verifying its own work and at managing long context."
The numbers the company published, always against Grok 4.6:
| Benchmark | Grok 4.7 | Grok 4.6 |
|---|---|---|
| CursorBench 4.0 | 46.3% | 40.4% |
| DeepSWE v1.1 | 71.0% | 65.2% |
| EEBench | 64.0% | 53.0% |
| Terminal-Bench 4.0 | 38.0% | 20.3% |
| Harvey Legal Agent | 19.6% | 15.8% |
| HealthBench Professional | 56.7% | 48.5% |
A real generation-over-generation jump. Terminal-Bench nearly doubled.
Now the detail that slipped past practically all of the coverage: the price didn't change. $2 input and $6 output per million tokens are exactly the same numbers as Grok 4.6, published side by side in the docs. No cut, no launch promo. You got performance for free, which is good — but don't confuse that with "it got cheaper," which is what half the headlines implied. (OpenRouter lists the model at $1.60/$4.80, below the official price; whether that's router policy or a promotion, xAI doesn't say.)
And a footnote worth noting because nobody writing in Portuguese noted it: the announcement, the documentation, and the OpenRouter catalog no longer say "xAI." They say SpaceXAI LLC. The company started signing with the new name alongside this launch.
Swapping the model in your product because a new number hit the timeline is the same mistake as picking a database off a blog benchmark: the decision isn't in the vendor's table, it's in running your task on both and measuring cost per completed task. That kind of measurement, done in front of everyone with the number showing up on screen, is what happens live every week in the Clã Beer and Code.
The number that changes the decision most: 0.85 seconds
This is the data point that repositions the entire model.
Measured by Artificial Analysis, time to first token:
| Model | Time to first token |
|---|---|
| Grok 4.7 (xhigh) | 0.85 s |
| GPT-6 Astra (max) | 259.33 s |
| Claude Fable 5.1 | 298.45 s |
That's not a typo. It's four to five minutes.
And it's not network slowness. It's reasoning. Models like Fable 5.1 and GPT-6 Astra, at their maximum effort settings, think before emitting the first character. For batch work, that's irrelevant: you fire it off and come back later. For anything with a human waiting in front of a screen, it's the difference between a product and a report.
Grok 4.7 answers in under a second.
That's what defines where it lives.
Where Grok 4.7 pays off
1. Anything on the synchronous request path
Chat inside your product. An assistant that answers while the user types. Ticket triage before routing. Intent classification. Field extraction from a form. Real-time suggestions.
All of that territory has a latency budget measured in seconds, not minutes. There, Grok 4.7's intelligence index of 46 (ranked 16th out of 202 models evaluated) is comfortably enough, and the 0.85 seconds are decisive. Putting a max-reasoning model on that path isn't over-engineering — it's breaking the product.
2. Native X search: nobody else has this
The Grok 4.7 tool list in the official docs includes function calling, web search, code execution, and X search. Searching X as a native tool of the model, without you building a scraper, without a side API, without an ingestion pipeline.
If your product touches brand monitoring, sentiment analysis on a live event, trend detection, due diligence on a public figure, or anything where the signal is born on X before it becomes news — this is a functional differentiator, not a benchmark number. No model from OpenAI, Anthropic, or Google has native access to that data.
In practice, it's the most defensible use case the model has today.
3. 500k context with cache at $0.50
Half a million tokens of window, with cached input at $0.50 per million — a quarter of the regular input price.
That works well for the fat, stable prompt pattern: a big system prompt full of business rules, a fixed document base, long few-shot. You pay full price once and cheap on every call after. For an internal assistant that loads the same 200-page manual into every conversation, the math works out nicely.
4. Reasoning as a dial, not a switch
The model accepts four effort levels: low, medium, high (default), and xhigh.
This matters more than it looks, because every bad number you're about to see in the next section was measured at xhigh. If your task doesn't need deep reasoning — and most product tasks don't — dropping to low or medium cuts output tokens, cuts cost, and cuts latency all at once. The same model becomes a different, economical product.
Where it doesn't pay off
Long-running coding agents
Here the number is uncomfortable, and it comes in two versions.
| Terminal-Bench 4.0 | Score | Source |
|---|---|---|
| Grok 4.7 | 38.0% | xAI's own measurement |
| Grok 4.7 | 26% | independent measurement |
| DeepSeek V4.1 Flash | 27% | independent measurement |
| Claude Fable 5.1 | 55% | independent measurement |
| GPT-6 Astra | 60% | independent measurement |
The 12-point gap between xAI's measurement and the third-party one isn't necessarily bad faith. Terminal-Bench is an agentic benchmark, and agentic benchmarks depend on the scaffold: the same model lands in a different tier depending on the harness wrapped around it. xAI even says in the announcement that it trained the model to "natively understand the Grok Bot harness" — which explains a good chunk of the gap and, at the same time, warns you that the 38% number may not follow you outside their environment. We already picked this effect apart in the DeepSeek V4.1 Flash comparison, where swapping the scaffold moved the result more than swapping the model.
But under both readings the practical conclusion is the same: on long agentic coding tasks, Grok 4.7 sits between 17 and 22 points behind the leaders. By the independent measurement, it ties with a Chinese flash model that costs a fraction of the price.
If your use case is an agent running for hours on a large repository, it's not the pick.
Output speed, if the text is long
39.5 tokens per second at xhigh, versus 65.8 for Fable 5.1 and 65.3 for GPT-6 Astra.
In other words: it starts fast and keeps going slow. For short responses that's great, because what the user perceives is the first token. For long text generation, the initial latency gain gets eaten up along the way.
Prompts above 200k tokens
The pricing isn't linear. Go past 200k prompt tokens and everything doubles:
| Prompt tier | Input | Cache | Output |
|---|---|---|---|
| Below 200k | $2.00 | $0.50 | $6.00 |
| 200k and up | $4.00 | $1.00 | $12.00 |
If your usage pattern lives in the upper half of the 500k window, you're not paying $2 per million. You're paying $4, with output at $12 — and at that point the price advantage over the premium models shrinks a lot.
Do not just follow the news — master it. Hands-on AI Engineering, live, every week, in the largest community in Brazil.
Join the ClãThe per-token pricing trap
Now the math that settles the argument. Besides list price, Artificial Analysis publishes the intelligence index's cost per task: how much it costs, in dollars, for the model to complete the same set of evaluations.
| Grok 4.7 (xhigh) | GPT-6 Astra (max) | Claude Fable 5.1 | |
|---|---|---|---|
| Intelligence index | 46 (#16/202) | 53 (#3/202) | 53 (#1/202) |
| Input / output price | $2 / $6 | $10 / $50 | $10 / $50 |
| First token | 0.85 s | 259.33 s | 298.45 s |
| Output | 39.5 tok/s | 65.3 tok/s | 65.8 tok/s |
| Cost per task | $3.74 | $3.26 | $7.63 |
Read the last row slowly.
Grok 4.7 is five times cheaper per token than GPT-6 Astra. And it comes out more expensive than Astra on the same battery of tasks: $3.74 versus $3.26.
The reason is mechanical, not mysterious. The model makes up for the capability gap by spending tokens: it reasons more, retries more, verifies more. Five times cheaper per unit doesn't help when you burn six times more units to get to the same place — and a slightly worse place, since the index is 46 versus 53.
Against Fable 5.1 the story flips: there Grok saves you half ($3.74 versus $7.63). So it's not that it's expensive. It's that "cheap per token" and "cheap per task" are two different metrics, and only the second one shows up on your invoice.
The rule that comes out of this is boring and applies to any new model: list price is a premise, not a conclusion. Run your task, count the tokens spent until you get an accepted answer, divide. It's the only measurement that has anything to do with your product.
Verdict by scenario
| Scenario | Pick | Why |
|---|---|---|
| Chat/assistant inside the product, user waiting | Grok 4.7 (low or medium) |
0.85 s to first token; an index of 46 is more than enough for this task |
| Triage, classification, extraction on the request path | Grok 4.7 (low) |
same reason, with minimal output cost |
| Product that reads signal from X in real time | Grok 4.7 | native X search; there's no equivalent |
| Assistant with a huge fixed prompt and lots of calls | Grok 4.7 | cached input at $0.50 with a 500k window |
| Coding agent running for hours on a large repo | GPT-6 Astra or Fable 5.1 | 60% and 55% on Terminal-Bench versus 26–38% |
| Hard reasoning, batch, latency doesn't matter | GPT-6 Astra | index of 53 and the lowest cost per task ($3.26) |
| High volume of simple tasks, cost above everything | flash/cheap model | Grok 4.7 isn't a price-floor model |
The one-sentence summary: Grok 4.7 is a latency model, not a frontier model. It competes where the clock is in charge, not where difficulty is.
Limitations and things to watch
- The announcement's benchmarks compare 4.7 only against 4.6. There's no official table against Fable 5.1, GPT-6 Astra, or Gemini. Every cross-model comparison in this post comes from third-party measurement, with a different scaffold than xAI's.
- The speed and cost-per-task data are from the
xhighprofile. Atlowormediumthe behavior changes a lot — probably for the better on cost and latency, for the worse on quality. If you're going to use those levels, measure, don't extrapolate. - The Grok 4.7 Fast variant exists, costs twice as much, and is restricted to Cursor and Grok Build. It's not on the open API.
- Knowledge cutoff of May 2026. For anything after that, you depend on the search tools — which is exactly where native X search tips the scale.
- X search is a feature, not a guarantee. X content is a noisy primary source. If your product is going to cite it to the end user, a verification layer isn't optional.
Quick FAQ
Is Grok 4.7 free?
On the API, no: it's billed from the first token, $2 per million input and $6 output. Grok Build has a free tier, but xAI's own pricing docs note that this tier doesn't include the Grok 4.7 Fast variant. For product use, treat the model as paid.
How much does Grok 4.7 cost?
$2 input and $6 output per million tokens, with cached input at $0.50. Past 200k prompt tokens, everything doubles: $4, $1, and $12. These are exactly the same numbers as Grok 4.6 — the price didn't drop this generation, only performance went up.
How do I access Grok 4.7?
The xAI API with the model id grok-4.7, plus Cursor, Grok Build, routers like OpenRouter, and third-party coding harnesses. It accepts text and image in, text out, with four reasoning effort levels: low, medium, high, and xhigh.
What does Grok 4.7 do differently from other models?
Two concrete things. It returns the first token in 0.85 seconds, versus minutes for the premium rivals at max reasoning. And it ships X search as a native tool — access that no model from OpenAI, Anthropic, or Google has.
Is it worth switching my coding agent to it?
If the agent runs long, autonomous tasks, no. It's 26% to 38% on Terminal-Bench 4.0 against 60% for GPT-6 Astra. If your "agent" is an interactive assistant that answers on the spot, then it's worth testing — the latency profile is a different game. The scenario-by-scenario criteria are in the coding tools comparison.
Why did xAI become SpaceXAI?
The "SpaceXAI LLC" signature appears in the announcement, in the documentation, and in the OpenRouter catalog starting with this launch. The company didn't publish a note explaining the corporate change alongside the model — what can be said today is that the name changed, not what motivated it.
What holds up
Grok 4.7 is a good model being sold on the wrong metric.
xAI positioned it as "the most capable for code and knowledge work," and the coding benchmarks are exactly where it falls furthest behind. Meanwhile, the number that actually sets the model apart — first token in 0.85 seconds, versus four minutes for the premium rivals — doesn't show up in any headline, because latency doesn't make for a pretty bar chart.
For anyone building product, it's the opposite: latency is what the user feels, and raw capability is what the engineer admires. The two rarely live in the same model, and picking the wrong one costs you retention.
The bigger lesson isn't about Grok. It's that price per token has become a marketing metric. $2 versus $10 looks like an obvious call until you measure cost per task and find out the final bill flips it. That holds for the next cheap model that gets announced, and there will be one in the coming weeks.
If you want to see the same effect with a different cast, the GPT-6 Astra, Fable 5.1, and GPT-5.6 Sol comparison has the cache math that makes an agent session 54% more expensive without anyone noticing.
For the record: published on September 22, 2026, one day after the launch. Third-party benchmarks on a freshly released model shift in the first few weeks — this is a living post and will be updated as new independent measurements come in.
{AI Engineer} — apaixonado por Laravel, arquitetura de software e construir produtos com impacto. Compartilho aqui tutoriais, descobertas e reflexões sobre o dia a dia de engenharia.
There is no shortage of content. What is missing is someone to untangle it: what matters now is how to implement it the right way. In the Clã you get that live, every week, with people who have already filtered out the noise.
Join the Clã