~/beer-and-code
▪ next event Workshop: Jev na Prática para Devs · 14 Oct · 19h — duração de 2h a 3h, ao vivo via Google Meet save your seat ›
~ / news / deepseek-v4-flash-0731-benchmarks-pricing $
News

DeepSeek V4 Flash 0731: New Weights, 82.7 on Terminal Bench and the Bill That Hurts OpenAI

LS Lucas Souza · · 12 min read
DeepSeek V4 Flash 0731: New Weights, 82.7 on Terminal Bench and the Bill That Hurts OpenAI

DeepSeek swapped the engine without changing the license plate. DeepSeek V4 Flash got new weights today, and if your agent points at deepseek-v4-flash, it woke up running a different model — with no changelog at all.

On July 31, 2026, DeepSeek put V4-Flash into public beta as build 0731. Same architecture, same size, same model name in the API. What changed was the post-training: the weights are different, and the agent benchmarks jumped in a way you can't ignore — DeepSWE went from 7.3 to 54.4.

In this post: what exactly changed, the numbers DeepSeek published versus the ones an independent shop measured, what this really costs against GPT-5.6 Luna after OpenAI's 80% cut, and what to do if you have an agent in production pointing at a model name that just became a moving target.

TL;DR

  • What it is: DeepSeek-V4-Flash-0731, the official release (public beta) of V4-Flash. Same architecture and size as the Preview, just re-post-trained.
  • Specs: 284B total parameters, 13B active, MoE, 1M context, 384K max output, MIT license.
  • Headline benchmark: 82.7 on Terminal Bench 2.1 (DeepSeek's number). Artificial Analysis measured 79% in its own test.
  • Price: $0.0028 input on cache hit, $0.14 input on cache miss, $0.28 output per 1M tokens.
  • The catch: the model name in the API didn't change. If you were calling deepseek-v4-flash yesterday, you're calling a different set of weights today.
  • Running it locally: Unsloth's Q8 quant is 161.9 GB and wants 169 GB of combined RAM+VRAM.

Update (Aug 12): the rumor became fact. The official V4 Pro shipped as DeepSeek-V4-Pro-0813, with a reported 87.9 on Terminal Bench and its own pricing table. The numbers, the Pro 0813 vs Flash 0731 comparison and the price increase warning are in DeepSeek V4 Pro 0813: 87.9 on Terminal Bench beats Opus 4.8.

What DeepSeek changed in V4 Flash 0731 (and what it didn't)

Start with what didn't change, because that's the confusing part.

The architecture is the same: 284B total parameters with 13B active per token, a 43-layer MoE backbone with 256 routed experts plus one shared expert per block, top-6 routing, 1M context. The DeepSeek documentation is explicit: 0731 "keeps the same architecture and size as DeepSeek-V4-Flash-Preview, and was only re-post-trained."

And the name in the API didn't change either. You keep sending model: "deepseek-v4-flash". Same string, different weights.

What changed was the post-training, and it changed a lot. Looking at the table in the Hugging Face card itself, the jump on agent tasks is the kind of thing that normally takes a new model generation:

Benchmark V4-Flash-0731 V4-Flash (Preview) V4-Pro (Preview)
Terminal Bench 2.1 82.7 61.8 72.1
DeepSWE 54.4 7.3 12.8
Cybergym 76.7 38.7 52.7
Toolathlon-Verified 70.3 49.7 55.9
DSBench-FullStack 68.7 37.0 41.8

Look at the V4-Pro column. Flash, the small cheap model, passed its big sibling — the Pro Preview has 1.6 trillion total parameters and 49B active. A model with 13B active beating one with 49B active on nine agent benchmarks says less about size and more about how much fat tool-specific post-training still has left to burn.

Here's the operational detail the card doesn't shout about: with no versioning in the model name, this swap landed in your production without warning. Keeping up with this on your own, every week, is what wears you down — and it's not a lack of discipline on your part, it's the lack of an environment. In the Clã Beer and Code this kind of read is routine: live execution every week, with people testing new models on real products before swapping out what's running in production.

DeepSeek V4 Flash benchmarks: wins in the terminal, loses everywhere else

Now the part the headline doesn't tell you.

The 82.7 on Terminal Bench is real and it's published. But it's not first place. In DeepSeek's own table, Opus 4.8 scores 85.0 on the same test. GLM-5.2 scores 81.0. In other words: V4-Flash-0731 joins the top-tier fight in the terminal, edges out GLM by a hair and stays behind Opus.

And once you leave the terminal, the picture gets worse:

Benchmark V4-Flash-0731 GLM-5.2 Opus-4.8
Terminal Bench 2.1 82.7 81.0 85.0
DeepSWE 54.4 46.2 58.0
NL2Repo 54.2 48.9 69.7
DSBench-Hard 59.6 54.5 71.7
Cybergym 76.7 — 83.1

A 15-point gap on NL2Repo and a 12-point gap on DSBench-Hard are not a detail. Translated to your day-to-day: the model got very good at operating a terminal and chaining tools, and it's still mediocre at whole-repository engineering tasks and hard data analysis.

There's also the outside data. Artificial Analysis put V4 Flash 0731 at 50 on the Intelligence Index — 10 points above the previous Flash, 6 points above V4 Pro (44), and one point behind GPT-5.6 Luna (max) and GLM-5.2 (max), both at 51. In their own terminal test, they measured 79%, not 82.7. That's not a contradiction: a benchmark run by the lab and a benchmark run by a third party almost never match, because the scaffold, the number of attempts and the harness all change. That's why the vendor's number works as a signal, not a contract.

And there's the cost hidden in tokens. Artificial Analysis logged ~206M output tokens just to run the index, and the Hacker News discussion (467 points, 254 comments) hammered exactly that: commenters reported the model needs a lot more tokens than a Gemini 3.6 Flash to close out the same task. A low per-token price with high verbosity is not the same thing as a low bill.

Honest summary: it's not a clean win. It's a cheap model that became competitive as a terminal agent, is in a technical tie on the general index and is still behind on repository SWE.

DeepSeek V4 Flash pricing: where it hurts OpenAI

DeepSeek's table, published in the official docs, per 1M tokens:

Price
Input (cache hit) $0.0028
Input (cache miss) $0.14
Output $0.28

The $0.0028 cache hit is the number that changes architecture, not spreadsheets. It's 50x cheaper than a cache miss. If your agent has a large, stable system prompt — and a terminal agent always does — the difference between structuring the prompt to hit cache and not structuring it is the difference between one bill and fifty.

Against OpenAI: Artificial Analysis calculated that the cost per task of V4 Flash 0731 on DeepSeek's first-party API comes out ~60% cheaper than GPT-5.6 Luna (max) — and that's already counting the 80% cut OpenAI applied to Luna. That's the part that hurts: OpenAI cut 80% and is still more expensive, against a model that ties it on the general index within one point.

Except "cheaper per task" carries the caveat from the previous section. If your use case is one of those where the model gets verbose, part of the savings sneaks back in through the back door as output tokens. The bill that matters is yours, with your prompt — not the index average. Also worth remembering that DeepSeek has already experimented with variable pricing by peak hours, so this shop's pricing table is something you check, not something you memorize.

▪ Clã Beer and Code

Do not just follow the news — master it. Hands-on AI Engineering, live, every week, in the largest community in Brazil.

Join the Clã

Running it locally: the weights are on Hugging Face, but they want 169 GB

MIT weights on Hugging Face doesn't mean it runs on your machine.

Unsloth's quants, documented in their guide:

  • UD-Q8_K_XL (lossless): 161.9 GB file, 169 GB of combined RAM + VRAM recommended.
  • UD-Q4_K_XL (near lossless): 155.1 GB file, 162 GB recommended.
  • UD-IQ3_XXS: 110–135 GB, the most realistic option for a mortal's machine.

Note that Q4 and Q8 differ by only 7 GB — the fine-grained MoE architecture doesn't compress linearly the way a dense model does, so dropping bits buys you less than you'd expect. On Hacker News, a user with a 128 GB MacBook Pro reported 20–25 tokens/s of generation with prefill at 400–450 tps. Usable for async tasks, painful for an interactive agent loop.

If your plan was a single GPU, we already ran those numbers in detail in the post on running V4 Flash locally on an RTX 5090, and the pattern repeats here: the bottleneck is never whether the model fits, it's total system memory. Same plot as Kimi K3 when it released its weights — open weights are licensing freedom, not hardware freedom.

One limitation that saves you time before you download 162 GB: V4 Flash is not multimodal. If your pipeline reads screenshots, scanned PDFs or images, it's out for that step.

Tracker: the official DeepSeek V4 Pro

(Update above: the official V4 Pro already shipped as DeepSeek-V4-Pro-0813 — see the link at the top of the post. The text below is what held on July 31.)

Clear label: RUMOR.

What's fact: DeepSeek-V4-Pro has existed since April 2026, with 1.6 trillion total parameters and 49B active — and it's published as a preview, not as an official release. It never left the preview label.

What's rumor: that the official version of V4-Pro ships "soon." That's circulating in r/LocalLLaMA threads and has no announcement, date or official changelog. Don't plan a migration on top of it.

What you can safely read from today's move: DeepSeek just showed it can pull a huge agent jump out of post-training alone, without touching the architecture. Applying the same recipe to Pro is the natural next step — and the current situation, with cheap Flash passing expensive Pro on nine agent benchmarks, is embarrassing enough to force their hand. But "natural" and "announced" are different things.

I'll update this section when a formal announcement lands.

What to do if your agent points at deepseek-v4-flash

The actionable part. A model name with no version is a moving target, and you have three concrete problems.

1. You don't know which version answered. Log the entire response, not just the content. The model field in the response is the only witness you'll have when behavior changes:

resp = client.chat.completions.create(
    model="deepseek-v4-flash",
    messages=msgs,
)

logger.info("llm_call", extra={
    "model_requested": "deepseek-v4-flash",
    "model_returned": resp.model,          # guarde SEMPRE
    "prompt_tokens": resp.usage.prompt_tokens,
    "cached_tokens": resp.usage.prompt_tokens_details.cached_tokens,
    "completion_tokens": resp.usage.completion_tokens,
})

cached_tokens is there on purpose: it's what tells you whether you're paying $0.0028 or $0.14. A low cache hit rate is the most common cost bug and the quietest one.

2. You have no way to pin the version on the first-party API. DeepSeek doesn't expose a dated slug — deepseek-v4-flash is all there is. If version pinning is a requirement for your product, the route is a third-party provider: OpenRouter publishes the build as a dated slug (deepseek/deepseek-v4-flash-0731), which gives you reproducibility in exchange for one more layer and slightly worse pricing. It's a trade-off, not a magic fix.

3. You don't know if it got better for your case. An agent benchmark going up 47 points doesn't guarantee your specific flow improved — heavy tool-focused post-training tends to mess with output format, verbosity and instruction following. Run your eval suite against the model today and compare with the baseline. If you don't have a saved baseline, that's the lesson of the week: without your own versioned evals, a silent weight update is undetectable until the customer complains.

Quick FAQ

Do I need to change my code to use 0731? No. The model name is still deepseek-v4-flash and the call format is the same. 0731 is already answering. The work isn't integration, it's validation — running evals and checking whether the new behavior works for your case.

Is it better than GPT-5.6 Luna? Depends on the task. On Artificial Analysis's Intelligence Index they're practically tied (50 versus 51). As a terminal agent DeepSeek is strong. On cost per task it comes out ~60% cheaper even after OpenAI's cut. On large-repository tasks and on multimodal, Luna takes it.

Can I use it commercially? Yes. The weights are under the MIT license on Hugging Face, which allows commercial use, modification and redistribution. The relevant caveat isn't licensing, it's customer compliance and the geopolitical risk we broke down in the post on sanctions on Chinese open source AI.

Can I run it on a 64 GB setup? Not on the good quants. The realistic floor is IQ3_XXS at 110–135 GB. Below that you're into aggressive quantization that degrades exactly the agent capability that makes this model worth it.

Conclusion

What happened on July 31 is small and the implication is big. DeepSeek re-post-trained V4-Flash, published new weights under the same API name, and delivered a model that wins in the terminal, ties the general index within one point and loses on repository SWE. Genuinely cheap, genuinely competitive, and far from the clean win the headline suggests.

The lesson that outlives the news cycle is a different one, and it's not about DeepSeek: you're building a product on top of a dependency that updates itself, with no version and no warning. That's not a Chinese lab defect — it's the nature of consuming a model as a service. If you have versioned evals and a log of the returned model, you find the change on the chart. If you don't, you find it in the support ticket.

The likely next chapter is V4-Pro getting the same post-training treatment. When it ships, the question that will matter isn't what it scored on Terminal Bench. It's whether you have a way to measure what it did to your product.

Lucas Souza
Written by
Lucas Souza

{AI Engineer} — apaixonado por Laravel, arquitetura de software e construir produtos com impacto. Compartilho aqui tutoriais, descobertas e reflexões sobre o dia a dia de engenharia.

▪ Clã Beer and Code

There is no shortage of content. What is missing is someone to untangle it: what matters now is how to implement it the right way. In the Clã you get that live, every week, with people who have already filtered out the noise.

Join the Clã
Meet the Clã Beer and Code
playing