DeepSeek V4.1 Flash vs Opus 5, Sol, and K3: 41x Cheaper and the Benchmark Nobody Quotes
DeepSeek dropped V4.1 Flash today and the timeline did what it always does: benchmark screenshot, red arrow on the number, "beat Opus 5 at 1% of the price". The price part is true; the "beat" part depends on which benchmark you look at.
DeepSeek V4.1 Flash is a multimodal MoE with a 552B-parameter backbone that activates 8B per token on prefill and 16B on decode, runs a 1M context window, and shipped under an MIT license. In DeepSeek's own official table it leads Terminal-Bench 2.1 and DeepSWE v1.1 — ahead of Opus 5, GPT-5.6 Sol, Kimi K3, and GLM-5.3. In that same official table, it gets crushed on Terminal-Bench 3.0 and 4.0.
Nobody is quoting those last two rows. This post does. We'll compare V4.1 Flash against its real market equivalents on price and benchmarks, show where the savings pay off and where they get expensive, and close with the number that surprised me most: the scaffold you use changes the result more than swapping the model does.
TL;DR
- What it is: DeepSeek V4.1 Flash, a 552B multimodal MoE (8B/16B active), 1M context, native vision, MIT license.
- Models compared: Claude Opus 5, GPT-5.6 Sol, Kimi K3, GLM-5.3, DeepSeek V4 Pro, GPT-6 Astra, Grok 4.6, Qwen 3.8 Flash.
- Cost/Access: $0.15 per 1M input tokens and $0.60 per 1M output tokens off-peak. Open weights on Hugging Face.
- The detail that changes the decision: it leads Terminal-Bench 2.1 with 90.6 and drops to 30.0 on 3.0 and 31.2 on 4.0, while Opus 5 holds 43.3 and 51.8.
- Useful links: official API changelog and model card on Hugging Face.
What DeepSeek shipped (and what dies on the 14th)
Let's start with the documented facts, because some of this is going to break code in your hands.
The new model goes by deepseek-flash in the API. The names deepseek-v4-flash and deepseek-v4-flash-vision-exp have been deprecated and are temporarily redirected to V4.1 Flash. And here's the notice that matters if you have an agent in production, from the official changelog:
Starting at 04:00 UTC on Sept 14, 2026, all
deepseek-v4-prorequests will route to V4.1-Flash at V4.1-Flash rates.
Translation: at 1 a.m. on Sunday, September 14, Brasília time, anyone calling deepseek-v4-pro wakes up running a different model. Different architecture, different behavior, different price. If that sounds familiar, it's because this is the second time: DeepSeek already swapped the V4 Flash weights without changing the API name back in July, and the V4 Pro 0813 we covered last month is exactly the model being retired now.
The engineering is genuinely interesting. The architecture is called Causal Encoder-Decoder: 40 layers arranged as 20 causal encoder layers followed by 20 decoder layers, with the decoder's global KV cache projected from the encoder's final hidden states instead of derived layer by layer. That's what makes it possible to activate only 8B parameters per token on prefill. Add Compressed Sparse Attention 2 and an FP4 KV cache, and the global cache footprint drops to 890 bytes per token — about 1/4 of the previous V4 Flash. For agentic workloads, which are input-heavy by nature, that's the whole game.
Picking between these models is the easy part of the decision. The hard part is building the harness and the eval that tell you whether the switch paid off on your code — and that's exactly what we build live, with errors showing up and getting debugged on the spot, at AI Engineering Lab 3ª Edição, on September 19 and 20.
Price: DeepSeek V4.1 Flash against its equivalents
This one is short, because the numbers leave no wiggle room. Prices in dollars per 1M tokens, pulled from OpenRouter on Sept 10, 2026:
| Model | Input | Output | Context |
|---|---|---|---|
| DeepSeek V4.1 Flash | 0.15 | 0.60 | 1M |
| Qwen 3.8 Flash | 0.15 | 0.47 | 1M |
| GLM-5.3 Flash | 0.15 | 0.50 | 1.3M |
| GPT-5.6 Luna | 0.20 | 1.20 | 1.05M |
| DeepSeek V4 Pro | 0.87 | 1.74 | 1M |
| GLM-5.3 | 1.40 | 4.40 | 1.3M |
| Grok 4.6 | 2.00 | 6.00 | 500k |
| GPT-5.6 Sol | 2.00 | 10.00 | 1.05M |
| Kimi K3 | 3.00 | 15.00 | 1M |
| Claude Opus 5 | 5.00 | 25.00 | 1M |
| GPT-6 Astra | 10.00 | 50.00 | 1.05M |
An Opus 5 output token costs 41 times what a V4.1 Flash one does. GPT-6 Astra's, 83 times. And notice something most comparison posts get wrong: V4.1 Flash is not the cheapest model in the table. Qwen 3.8 Flash and GLM-5.3 Flash charge the same on input and less on output. The $0.15 tier already has three contenders — what V4.1 Flash brings that's new isn't the price, it's what it delivers at that price.
Keep in mind that direct pricing on DeepSeek's API has peak-hour rates: the $0.15 and $0.60 are off-peak. Between 01:00–04:00 and 06:00–10:00 UTC, Monday through Friday, it doubles.
Benchmarks: where it actually wins
This is the official table from the model card, with reasoning effort maxed out. It's not a tweet screenshot, it's the document DeepSeek itself published:
| Benchmark | Opus 5 | GPT-5.6 Sol | Kimi K3 | GLM-5.3 | V4 Pro | V4.1 Flash |
|---|---|---|---|---|---|---|
| GPQA Diamond | 93.4 | 94.1 | 92.9 | 88.1 | 92.4 | 90.9 |
| HLE | 56.3 | 44.5 | 43.5 | 42.0 | 42.7 | 36.8 |
| Terminal-Bench 2.1 | 89.1 | 88.8 | 88.3 | 88.2 | 87.9 | 90.6 |
| DeepSWE v1.1 | 74.0 | 73.0 | 67.5 | 66.9 | 62.7 | 74.2 |
| Codeforces (rating) | — | — | — | — | 3348 | 3471 |
| CyberGym | — | 84.5 | 80.0 | 84.5 | 83.3 | 88.1 |
| AutomationBench | 50.3 | 45.8 | 46.7 | 48.8 | 43.2 | 54.8 |
| Agent's Last Exam | 28.6 | 26.7 | 27.6 | 28.5 | 25.7 | 31.8 |
| NL2Repo-Bench | 75.3 | 56.8 | 58.0 | 58.0 | 61.5 | 64.0 |
| ProgramBench | 37.0 | 23.0 | 17.5 | 19.0 | 15.5 | 20.3 |
Read it slowly, because there are two stories here.
The first one is real and it's good: on agentic terminal tasks and issue resolution — Terminal-Bench 2.1, DeepSWE v1.1, AutomationBench, CyberGym — V4.1 Flash leads. It doesn't tie: it leads, ahead of models that cost 40 times more. If you run agents that open PRs, fix reported bugs, and work inside existing repositories, that's money in your pocket.
The second one is the counterweight: on pure reasoning and building from scratch, it gets beaten. HLE 36.8 against Opus 5's 56.3. ProgramBench 20.3 against 37.0. NL2Repo-Bench 64.0 against 75.3. GPQA Diamond below Opus 5, below Sol, and below the very V4 Pro it's replacing.
The pattern is consistent and it makes architectural sense: a model that activates 8B parameters per input token is excellent at navigating large context cheaply, and limited at synthesizing new, deep reasoning.
Do not just follow the news — master it. Hands-on AI Engineering, live, every week, in the largest community in Brazil.
Join the ClãThe benchmark nobody is quoting
Now for the part that vanished from every screenshot making the rounds today. Terminal-Bench doesn't have just one version:
| Opus 5 | GPT-5.6 Sol | GLM-5.3 | Kimi K3 | V4.1 Flash | |
|---|---|---|---|---|---|
| Terminal-Bench 2.1 | 89.1 | 88.8 | 88.2 | 88.3 | 90.6 |
| Terminal-Bench 3.0 | 43.3 | 34.4 | 28.3 | 17.7 | 30.0 |
| Terminal-Bench 4.0 | 51.8 | 39.9 | 37.9 | 12.6 | 31.2 |
2.1 is the benchmark everyone is quoting, and on it V4.1 Flash takes first place with 90.6.
On 3.0 it drops to 30.0 — below Opus 5 and GPT-5.6 Sol. On 4.0 it drops to 31.2, while Opus 5 climbs to 51.8. A gap of 20 absolute points.
Why does that happen? Because Terminal-Bench 2.1 is saturated. Five different models scoring between 87.9 and 90.6 isn't a benchmark measuring capability — it's a benchmark measuring who trained closest to its task set. Versions 3.0 and 4.0 exist precisely because 2.1 stopped discriminating. And that's where the distance between a frontier model and a cheap model shows up again in full.
This isn't an accusation of bad faith against DeepSeek — the 3.0 and 4.0 numbers are published in their official table, nothing hidden. The problem is the hype cycle, which crops out the good row and throws away the two below it.
And it lines up with the independent test that ran in parallel. Pawel Huryn tested the model on two repositories with 105 hidden bugs, asking it to find and fix whatever it could:
| Model (effort) | Bugs fixed | Cost |
|---|---|---|
| Opus 5 (max) | 27 | $51.33 |
| Grok 4.6 (max) | 27 | $16.96 |
| DeepSeek V4.1 Flash (max) | 24 | $1.80 |
| GPT-5.6 Luna (xhigh) | 23 | $2.50 |
| Opus 5 (high) | 21 | $38.77 |
That's not the "crushed Opus 5" scoreboard that went around. It's 24 against 27 — 11% fewer bugs fixed for 3.5% of the cost. Which, let's be honest, is still an absurd result. It's just not the result the screenshot promised.
Look at the Opus 5 (high) row too: 21 bugs for $38.77. The same model, with less reasoning effort, delivers less than V4.1 Flash while charging 21 times more. The wrong configuration costs more than the wrong model.
The scaffold matters more than the model
This is the number I most wish I'd seen a year ago, and it's buried at the end of the model card. Same model, same tasks, same sampling settings — changing only the harness that orchestrates the agent:
| Scaffold | DeepSWE v1.1 | Terminal-Bench 2.1 |
|---|---|---|
| mini-SWE | 74.2 | 90.3 |
| DSH Minimal | 72.6 | 90.6 |
| DSH Standard | 70.5 | 85.8 |
| Claude Code | 69.8 | 88.0 |
| DSH PTC | 67.6 | 85.8 |
| Pi | 66.2 | 86.1 |
| Codex | 65.6 | 84.1 |
| OpenCode | 65.5 | 85.0 |
An 8.7-point spread on DeepSWE, with the identical model. Just by swapping the scaffold.
Compare that with the gap that sparked today's entire debate: V4.1 Flash's 74.2 against Opus 5's 74.0 on the same benchmark. Two tenths.
In other words: the choice of harness weighed forty times more on the final result than the choice between the cheapest and the most expensive model on the market. If your agent is doing badly, the odds that the problem is in your orchestration loop — how you assemble context, how many steps you allow, how you handle tool output — are far higher than the odds that it's the model.
That's why "which model should I use" is almost always the wrong question. And that's why switching models rarely fixes what we expect it to fix.
Limitations and things to watch
The V4 Pro rerouting is on the 14th and it's not optional. If you have a production agent pointing at deepseek-v4-pro, you have four days to test V4.1 Flash in your workflow. Behavior will change — V4 Pro is a 1.6T MoE with 49B active, V4.1 Flash is 552B with 8B/16B. It's not the same brain.
Running it locally isn't for just any machine. MIT weights on Hugging Face doesn't mean it fits on your GPU. That's 552B of backbone plus 196B of Engram. The license is open; the VRAM bill, not so much.
Rates double at peak. Every savings calculation done with $0.15 and $0.60 assumes an off-peak window. If your workload runs during Brazilian business hours, check which UTC band it lands in before promising savings to your team.
Reasoning effort isn't a detail. The model accepts continuous effort from 1 to 100, and every number in this post is at 100. Running at low effort tanks results and cost together — and it's a variable you need to pin down before comparing anything with anything.
A saturated benchmark is noise. If a benchmark has five models separated by 2.7 points, it has stopped helping you decide. Look for the newer version, or build your own task set from the code you actually maintain.
Quick FAQ
Is it worth switching my agent from Opus 5 to V4.1 Flash today? If the work is agentic and repetitive — fixing reported bugs, working in an existing repo, running terminal commands, triage — it's worth testing right now: the benchmarks favor it and the savings are more than an order of magnitude. If the work is new architecture, long reasoning, or building from scratch, the NL2Repo-Bench and ProgramBench numbers say stay where you are.
Which model name do I call in the API?
deepseek-flash points to V4.1 Flash. deepseek-v4-flash and deepseek-v4-flash-vision-exp are temporarily redirected, and deepseek-v4-pro starts being redirected at 04:00 UTC on Sept 14, 2026. Pin the new name before that date.
Can I use it commercially? Yes. The weights are under an MIT license on Hugging Face. Through the API, DeepSeek's terms apply. If your context is a company with sensitive data, the conversation is about governance and data residency, not licensing — and that's where OpenRouter or a managed provider changes the math.
Is it truly multimodal or is it an adapter bolted on afterward? Truly multimodal. The vision encoder was trained from scratch and the visual embeddings go in alongside the text ones from the start of pretraining, on a 45T-token multimodal corpus. In practice: DocVQA 95.6 and RefCOCO 86.0 on the base model.
Conclusion
DeepSeek V4.1 Flash is an important release and the savings are real — for agentic workloads, it delivers near the top while charging a fraction. That's not hype, it's the official table.
But the post that only shows the Terminal-Bench 2.1 row is selling you half the information. The other half is that the same model falls 20 points below Opus 5 on the newer versions of the same benchmark, and that the scaffold you pick moves the result forty times more than swapping the model.
The lesson left over isn't about DeepSeek. It's about method: if you don't have a task set from your own code and a stable harness to run both options side by side, you're not choosing a model. You're choosing a headline.
If you want to see an agent's token bill built line by line, the next step is how much an AI agent costs in production. And to understand why frontier benchmarks keep drifting away from real-world experience, check out the GPT-6 Astra vs Fable 5.1 vs GPT-5.6 Sol comparison.
{AI Engineer} — apaixonado por Laravel, arquitetura de software e construir produtos com impacto. Compartilho aqui tutoriais, descobertas e reflexões sobre o dia a dia de engenharia.
There is no shortage of content. What is missing is someone to untangle it: what matters now is how to implement it the right way. In the Clã you get that live, every week, with people who have already filtered out the noise.
Join the Clã