DeepSeek V4 Pro 0813: 87.9 on Terminal Bench Beats Opus 4.8, and Your Bill Is Going Up
Twelve days ago DeepSeek's cheap model beat DeepSeek's expensive model on nine agent benchmarks. It was embarrassing. Today the flagship is back.
DeepSeek V4 Pro 0813 showed up on DeepSeek's official pricing page under the model ID DeepSeek-V4-Pro-0813, and it's already serving on the API. This is the official release of V4 Pro, the one that had been stuck with a preview label since April. DeepSeek's changelog, however, still doesn't have a single line about it.
In this post: what exactly went live today, which numbers are fact and which are reporting, how 0813 stacks up against Flash 0731 and Opus 4.8, and the notice written on DeepSeek's own pricing page that changes your bill over the next few weeks.
TL;DR
- What it is:
DeepSeek-V4-Pro-0813, the official (GA) release of DeepSeek V4 Pro. It had been in preview since April 2026. - Specs: MoE with 1.6 trillion total parameters and 49B active, 1M context, 384K max output, thinking and non-thinking modes, tool calling.
- Pricing: $0.003625 input on cache hit, $0.435 input on cache miss, $0.87 output per 1M tokens.
- Headline benchmark: 87.9 on Terminal Bench 2.1 — a reported number, not yet confirmed in the official changelog.
- The notice: DeepSeek's pricing page says, in its own words, that the company plans to raise API prices "in the near future," with a "significant increase expected."
- Where it already is: DeepSeek's first-party API and OpenRouter, with a dated slug.
What went live today (and what hasn't yet)
It's worth separating the two, because the internet is already treating all of it as an official announcement, and that's not quite what this is.
What's fact, with a primary source. The model ID DeepSeek-V4-Pro-0813 is published on DeepSeek's pricing page with its own table. OpenRouter already lists the model as a dated slug and describes it as "the GA release of DeepSeek V4 Pro". In other words: the model exists, has a price, is answering requests, and has dropped the preview label.
What isn't fact yet. DeepSeek's official changelog is still stuck on July 31, at the V4-Flash entry. And that entry ends exactly like this: "This update only upgrades the DeepSeek-V4-Flash API. The DeepSeek-V4-Pro API and the APP/WEB models are unchanged. The official release of DeepSeek-V4-Pro will follow soon."
"Soon" was today. Except the benchmark table that normally ships with a DeepSeek release isn't out yet, and the Hugging Face card for 0813 isn't public yet.
That matters for a practical reason: the numbers you're seeing passed around today didn't come from a DeepSeek document you can open. They came from people reading the pricing page and reporting the internal tables. I treat them as reporting in this post, and I flag each one.
Choosing between Pro 0813, Flash 0731, and Opus 4.8 this week isn't about reading a benchmark table — it's about deciding with incomplete information and having the range to know what the table doesn't tell you. That kind of reading is what we do together every week in Clã Beer and Code: it's paid, it's a subscription, and it's people testing new models on real products before swapping out what's already in production.
The DeepSeek V4 Pro 0813 numbers
Here's the consolidated table. The Pro Preview and Flash 0731 columns come from the official table on the V4-Flash-0731 card on Hugging Face — those are DeepSeek's own numbers. The Pro 0813 column is the reported one, with no public document yet.
| Benchmark | Pro 0813 (reported) | Pro Preview (official) | Flash 0731 (official) | Opus 4.8 (DeepSeek table) |
|---|---|---|---|---|
| Terminal Bench 2.1 | 87.9 | 72.1 | 82.7 | 85.0 |
| DeepSWE | 62.7 | 12.8 | 54.4 | 58.0 |
| CyberGym | 83.3 | 52.7 | 76.7 | 83.1 |
| DSBench-Hard | 67.2 | 31.1 | 59.6 | 71.7 |
| AutomationBench | 31.8 | 12.8 | — | — |
Three takeaways worth more than the raw numbers.
First: the Pro's jump belongs to the same family as the Flash's jump. The Flash had gone from 61.8 to 82.7 on Terminal Bench through post-training alone, without touching the architecture. The Pro now goes from 72.1 to 87.9. It's the same recipe applied to the big model — exactly what we bet would happen in the Flash 0731 post, when we marked the official V4 Pro as a rumor. It's now fact.
Second: DeepSWE going from 12.8 to 62.7 is the number that should be in the headline. Terminal Bench measures operating a terminal. DeepSWE measures solving real software engineering tasks, in a repository. A model going from 12.8 to 62.7 on the same architecture isn't a minor tweak, it's different behavior. If that number holds up, it's the difference between a model you use to run commands and a model you let touch code.
Third: 87.9 beats Opus 4.8 on DeepSeek's own yardstick. In the table DeepSeek published in July, Opus 4.8 scored 85.0 on Terminal Bench 2.1. If 87.9 is real, it's the first time an open-weights model shows up ahead of Opus on this test, as measured by the shop that built the model.
And that's where the usual caveat comes in: a benchmark run by the lab and a benchmark run by a third party almost never match, because scaffold, number of attempts, and harness all change. With Flash 0731, DeepSeek published 82.7 and Artificial Analysis measured 79% on its own test. Expect a similar pullback here when the independent measurement lands. A vendor number is a signal, not a contract.
DeepSeek V4 Pro pricing: the cache hit and the notice nobody is reading
The official table, per 1M tokens:
| V4 Pro 0813 | V4 Flash 0731 | |
|---|---|---|
| Input (cache hit) | $0.003625 | $0.0028 |
| Input (cache miss) | $0.435 | $0.14 |
| Output | $0.87 | $0.28 |
A cache hit on the Pro costs $0.003625. A cache miss costs $0.435. That's a 120x difference.
Let that sink in, because it's the most actionable piece of information in this entire post. If your agent has a large, stable system prompt — and terminal agents always do — the difference between structuring the prompt to hit cache and not structuring it isn't spreadsheet savings. It's an architecture decision. You pay the price of a cheap model or the price of an expensive model while running exactly the same model.
And you can measure that today, in one log line. The field that matters is cached_tokens:
resp = client.chat.completions.create(
model="deepseek-v4-pro",
messages=msgs,
)
usage = resp.usage
cached = usage.prompt_tokens_details.cached_tokens
hit_rate = cached / usage.prompt_tokens if usage.prompt_tokens else 0
logger.info("llm_call", extra={
"model_returned": resp.model, # qual build respondeu de verdade
"prompt_tokens": usage.prompt_tokens,
"cached_tokens": cached,
"cache_hit_rate": round(hit_rate, 3), # abaixo de 0.8 é dinheiro no lixo
})
A low cache hit rate is the most common cost bug and the quietest one: it doesn't break anything, it doesn't show up in tests, it just lands on the invoice. And model_returned in that same log is what will tell you, two weeks from now, whether DeepSeek swapped the weights again without telling anyone.
Now the notice. It's written on DeepSeek's own pricing page:
DeepSeek plans to raise overall pricing for DeepSeek API services in the near future, with a significant increase expected.
DeepSeek is warning that it's going to raise prices, and using the word "significant." There's no date and no percentage. But the message is clear: the table you're reading today is the cheapest table you'll ever see for this generation.
That changes what you do with this information. If you were waiting for things to "mature" before testing V4 Pro, the math has flipped: running your eval now, on the current table, is what gives you a baseline for comparison when the price goes up. Keep in mind this is a shop that has already experimented with variable pricing by peak hours — a DeepSeek price table is something you check, not something you memorize.
Do not just follow the news — master it. Hands-on AI Engineering, live, every week, in the largest community in Brazil.
Join the ClãDeepSeek V4 Pro vs Flash: Pro 0813 or Flash 0731?
The Flash didn't turn into garbage today. It just ended up with a narrower niche.
The Pro costs 3.1x the Flash on cache-miss input and 3.1x on output. In exchange, if the reported numbers hold up, it delivers ~5 more points on Terminal Bench, ~8 on DeepSWE, and ~7 on CyberGym. That's not a jump that justifies switching everything by default.
Where the Pro starts to pay off:
- Whole-repository tasks. DeepSWE at 62.7 versus the Flash's 54.4, and especially the Flash's track record of being weak on NL2Repo, point the Pro at this niche.
- Truly long context. Both have 1M, but 49B active parameters sustain long reasoning better than 13B. If your use case is analyzing an entire codebase, this is where it belongs.
- When mistakes are expensive. An agent loop that commits, touches infra, or touches customer data. Five benchmark points turn into many hours of debugging.
Where the Flash keeps winning:
- Volume. Classification, extraction, routing, summarization. Paying 3x for a task the Flash handles is burning budget.
- Latency. 13B active parameters respond faster than 49B. In an interactive loop that matters more than a benchmark.
- Running locally. The Flash already asks for 169 GB of RAM+VRAM at the good quant, as we detailed in the post on running V4 Flash locally. The Pro has 1.6 trillion total parameters. For a mere mortal's machine, this model is an API, not self-hosting.
Limitations and things to watch
The official card isn't out. None of the 0813 numbers has a public DeepSeek document behind it so far. If you're going to make an architecture decision on top of this today, make it knowing the source is reporting, and recheck when the changelog goes up.
The model name on the API is still a moving target. That was the central topic of the Flash 0731 post, and it hasn't changed: on the first-party API you call deepseek-v4-pro and there's no way to pin a build. If version pinning is a requirement for your product, the path is a third-party provider — OpenRouter publishes the dated slug deepseek/deepseek-v4-pro-0813, which gives you reproducibility in exchange for one more layer and slightly worse pricing.
An agent benchmark is not your product's benchmark. A 50-point jump on DeepSWE doesn't guarantee your flow got better. Heavy post-training on tools shifts output format, verbosity, and instruction adherence — things that can break your parser without showing up in any table.
Compliance before licensing. The V4 Pro weights are under MIT, which settles the legal side of commercial use. What a license doesn't settle is geopolitical risk and your customer's policy, which we broke down in the post on sanctions on Chinese open source AI.
Quick FAQ
Do I need to change my code to use 0813?
No. The name on the API is still deepseek-v4-pro. If you were already calling the Pro, 0813 is already answering you. The work isn't integration, it's validation: run your eval and compare against the baseline before assuming it got better.
How do I pin this specific version?
On DeepSeek's first-party API, you don't. Use OpenRouter's dated slug (deepseek/deepseek-v4-pro-0813) if reproducibility is a requirement. And always log the model field from the response — it's your only witness when behavior changes.
Can I use it commercially? Yes, the weights are MIT. The relevant caveat is customer compliance, not licensing.
Is it worth migrating from the Flash to the Pro right now? Only if your bottleneck is repository tasks or very long context. For volume and latency the Flash is still the right call. Migrate on measurement, not on headlines.
Conclusion
Today's fact is small: a new model ID on a pricing page. The implication is big.
DeepSeek has shown for the second time in twelve days that it can squeeze a big agent jump out of post-training alone, without touching the architecture. First on the Flash, now on the Pro. And 87.9 on Terminal Bench, if it holds, puts an open-weights model ahead of Opus 4.8 for the first time on DeepSeek's own yardstick.
But the part that outlives the news cycle is the notice on the pricing page. DeepSeek is signaling a significant increase. The cheap open model that was eating into the American labs' margins just announced it's going to get less cheap. If your architecture has a cost assumption baked into it — and nearly every agent architecture does — that assumption has an expiration date.
Run your eval while the table is still this one.
{AI Engineer} — apaixonado por Laravel, arquitetura de software e construir produtos com impacto. Compartilho aqui tutoriais, descobertas e reflexões sobre o dia a dia de engenharia.
There is no shortage of content. What is missing is someone to untangle it: what matters now is how to implement it the right way. In the Clã you get that live, every week, with people who have already filtered out the noise.
Join the Clã