~/beer-and-code
▪ next event Workshop: Jev na Prática para Devs · 14 Oct · 19h — duração de 2h a 3h, ao vivo via Google Meet save your seat ›
~ / news / gemini-3-7-flash-pricing-benchmarks $
News

Gemini 3.7 Flash Is Here: 43.6% on FrontierCode and $0.75/M, with Flash Ahead of Pro Again

LS Lucas Souza · · 11 min read
Gemini 3.7 Flash Is Here: 43.6% on FrontierCode and $0.75/M, with Flash Ahead of Pro Again

The discount has an expiration date, and nobody read the fine print

Google launched Gemini 3.7 Flash on August 13, 2026, described in the announcement as "our most intelligent workhorse model yet for coding and agents". The timeline filled up with screenshots of the same number: $0.75 per million input tokens. Cheap, right?

No. That's the hole in the coverage.

The official Gemini API pricing table lists 3.6 Flash and 3.7 Flash with identical values on every row: $1.50 input, $7.50 output. 3.7 isn't a cheaper model. It's the same price with a 50% discount through December 31, 2026. On January 1 your bill doubles without a single line of your code changing.

In this post: both prices and the date they flip, what 43.6% on FrontierCode is worth in production code, where the model already runs, what breaks when you migrate from 3.6, and why Google shipped its fifth Flash in a row before the Pro.

TL;DR

  • What it is: gemini-3.7-flash, Google's coding and agents model. A refinement on top of 3.6 Flash (the model card says "Based on Gemini 3.6 Flash"), not a new pretraining run.
  • Price: $0.75 / $3.75 per 1M through December 31, 2026. After that, $1.50 / $7.50 — what 3.6 Flash already charges.
  • Context: 1,048,576 input tokens and 65,536 output tokens, per the official model docs. The announcement didn't include that number; the docs do.
  • The real jump: FrontierCode 1.1 Main 43.6% vs 34.4%, DeepSWE v1.1 65.3% vs 49.0%, and AutomationBench 30.4% vs 17.0% against 3.6 (Google's numbers).
  • The catch: hallucination went up from 55.6% to 64.5% as measured by Artificial Analysis. It's right more often when it knows, and guesses more often when it doesn't.
  • Access: Gemini API, AI Studio, Android Studio, Antigravity, Gemini Enterprise, and Spark. Brazil included. No open weights, no self-hosting.

Gemini 3.7 Flash pricing: both numbers and the date it doubles

Here are the numbers, all from the official pricing page:

Row Through Dec 31, 2026 Starting Jan 1, 2027
Standard input $0.75 / 1M $1.50 / 1M
Output (including thinking) $3.75 / 1M $7.50 / 1M
Batch API $0.375 / $1.875 $0.75 / $3.75
Context caching $0.075 / 1M (+ $0.50/1M per hour of storage)

Look at the second row. The table literally says "Output price (including thinking tokens)". You pay $3.75 per million tokens the model thought and that never reach your application. A cost spreadsheet that only counts the returned text is wrong — and even more wrong at thinking_level: high, the level the docs describe as the one that maximizes the model's ability to think and use tools.

Simon Willison went straight to the point on Hacker News (956 points and 485 comments as I write this): the price doubles at the end of the year, and he asked who, exactly, plans to still be using this model five months from now. Fair question, considering 3.6 Flash was three weeks old when 3.7 shipped.

And that's where the engineering decision lives. Swapping models every week isn't a strategy. Strategy is having an eval set for your domain that tells you in ten minutes whether the new model improves your pipeline or just improves the vendor's slide — and that difference between using AI and building with AI is what we train every week in the Clã Beer and Code.

Gemini 3.7 Flash benchmarks: FrontierCode 43.6% vs 34.4% on production code

VERIFIED against the primary source: 43.6% on FrontierCode 1.1 Main versus 34.4% for 3.6 Flash. A caveat that applies to the entire release: FrontierCode, DeepSWE, GDP.pdf, and AutomationBench are Google's own benchmarks. Directional, not an independent audit.

What those points are worth in practice is in the behavior description, not the scoreboard. The announcement says the model "better adapts to roadblocks, clarifies intent when needed, and follows instructions with greater fidelity", and that "a more disciplined execution means less manual oversight and fewer retries". Translated for anyone maintaining an agentic pipeline: fewer loops hitting the same wall, fewer malformed tool calls, fewer retries. That matters more than the benchmark, because a retry is double the cost.

The strongest number in the release isn't even FrontierCode. It's AutomationBench: 30.4% versus 17.0% for 3.6 Flash. According to MarkTechPost's comparison, that puts 3.7 Flash above Claude Sonnet 5 (10.7%) and GPT-5.6 Terra (23.6%).

The overall scoreboard, though, is more modest. The Artificial Analysis Intelligence Index gives 3.7 Flash a 56, below GPT-5.6 Terra and Muse Spark 1.2 (57 each). Terra leads DeepSWE (69.6% versus 65.3%) and Terminal-bench 2.1 (87.4%). The case for Flash was never "it's the best". It's intelligence per dollar: at an 80/20 mix, the blended price comes out to $1.35/1M versus $3.60 for Sonnet 5 and $4.00 for Terra.

Where you can use it today

Broad rollout on day 1, according to the official announcement:

  • Gemini API / AI Studio — gemini-3.7-flash, stable version.
  • Android Studio, Google Antigravity (the agentic environment), and Gemini Enterprise.
  • Gemini Spark — Google AI Pro and Ultra subscribers in 160+ countries, excluding the European Economic Area, Nigeria, Switzerland, and the United Kingdom. Brazil is in.
  • OpenRouter, which listed the model that same August 13.

That last item isn't a footnote. The most upvoted criticism on HN wasn't about benchmarks: it was about the friction of getting a Google API key. jjcm summed it up as "routing you through 8 different dashboards to set up ACLs before you've hired your 2nd employee"; SyneRyder closed with "it was easier to go through OpenRouter than spend more energy on it". For a solo Brazilian dev with no GCP contract, that's the real pain.

▪ Clã Beer and Code

Do not just follow the news — master it. Hands-on AI Engineering, live, every week, in the largest community in Brazil.

Join the Clã

Gemini 3.7 Flash vs 3.6 Flash: is it worth switching in your pipeline?

Before answering: this is not a string swap. The official migration docs require you to remove the deprecated sampling parameters and replace the numeric thinking_budget with the thinking_level enum:

# Antes — 3.6 Flash e anteriores
resp = client.models.generate_content(
    model="gemini-3.6-flash",
    contents=prompt,
    config={
        "temperature": 0.2,
        "top_p": 0.95,
        "thinking_config": {"thinking_budget": 2048},  # numérico
    },
)

# Agora — 3.7 Flash
resp = client.models.generate_content(
    model="gemini-3.7-flash",
    contents=prompt,
    config={
        # temperature / top_p / top_k: removidos
        "thinking_config": {"thinking_level": "low"},  # low | medium | high
    },
)

A detail that kills an entire use case: minimal is no longer supported. The docs list "low, medium, high levels supported; minimal not supported". Anyone using the minimum floor for high-volume classification, routing, or extraction just lost the ground under their feet. That workload goes to Flash-Lite or out of Google entirely.

And before you celebrate the savings, measure what you actually pay:

u = resp.usage_metadata
billed_output = u.candidates_token_count + u.thoughts_token_count
custo = (u.prompt_token_count / 1e6) * 0.75 + (billed_output / 1e6) * 3.75
print(f"pensou {u.thoughts_token_count} / entregou {u.candidates_token_count} -> US$ {custo:.6f}")

Run it on your real cases, before and after. It's the only way to know whether the savings exist in your workload or only on the pricing card.

Short answer: it's worth testing if you run multi-step agentic tasks with tools. It's not worth switching blind if your pipeline is RAG, structured extraction, or anything where "I don't know" needs to be a valid answer — the why is in the next section. And if you didn't follow its predecessor, our coverage of Gemini 3.6 Flash with benchmarks and pricing is the baseline this release was built against.

Limitations and things to watch

This is the part the release doesn't tell you.

1. Hallucination went up along with accuracy. 64.5% versus 55.6% for 3.6 Flash on the AA-Omniscience measurement from Artificial Analysis, which penalizes answering wrong when the model should have declined. A regression for RAG and extraction. Deploying here without your own eval set is deploying in the dark.

2. Not every number went up. CharXiv Reasoning dropped from 85.2% to 84.5% without tools and from 89.4% to 88.7% with tools, per eesel's read of Google's table. Reasoning over charts got worse. Not by much, but it kills the "better at everything" claim.

3. Latency at high rules out interactive use. Artificial Analysis measured 9.83 seconds to first token. The 340.1 tokens/s is peak generation, not end-to-end experience. An IDE copilot at high is a user watching a cursor blink.

4. It's not the cheapest model, not even close. GPT-5.6 Luna costs $0.20 / $1.20 since July 30, 2026, after OpenAI cut the price by 80%. About 3.75x cheaper on input even against Flash's promo price, and ~7.5x after it ends.

5. No Live API, no open weights, no air-gapped deployment. Text-only output. Regulated industry or data sovereignty: it's out.

6. [UNCONFIRMED] Single-source reviews report that high generates ~40% more tokens than 3.6 on the same task and that real throughput via OpenRouter lands well below the peak. A signal to measure, not a fact. Same goes for the partner testimonials in the announcement (Browser Use talks about "35% cheaper than 3.6 Flash"): that's marketing, not a benchmark.

The Flash-before-Pro pattern and what it signals

Look at the sequence. Sundar Pichai promised Gemini 3.5 Pro at Google I/O on May 19, 2026. Google pushed it to July 17, citing a ground-up architectural redesign, and missed the deadline again. In the meantime we got 3.5 Flash, Flash-Lite, Flash Cyber, 3.6 Flash (July 21), and now 3.7 (August 13).

Five Flashes. Zero Pros.

The sober read is that Google ships in the tier where it's able to ship. Refining an existing model is a short cycle — hence the "Based on Gemini 3.6 Flash" in the model card. Training a new Pro is a long cycle, and the long cycle is stuck. If you've been following the saga, we have the Gemini 3.5 Pro launch tracker.

What does this signal about an eventual 3.7 Pro? Nothing guaranteed: any date is speculation. But the cadence suggests the Flash family remains Google's real product for as long as the Pro doesn't ship. Plan on replacement in weeks, not years.

Quick FAQ

Is the price really going to double on January 1? That's what the official table says: introductory through December 31, 2026, $1.50 / $7.50 after. Some people on HN are betting the discount gets extended for lack of buyers at full price, but that's speculation. Size your unit economics on the 2027 price.

Can I use it from Brazil? Yes. Brazil is among the 160+ countries in the Spark rollout, which excludes the EEA, Nigeria, Switzerland, and the United Kingdom. No restrictions have been announced for the API. The real friction is setting up billing on Google Cloud, and a lot of people get around it via OpenRouter.

Does it run locally? Are there open weights? No and no. API and enterprise only, inside Google's cloud. No self-hosting, no air-gapped deployment.

What's the knowledge cutoff? March 2026, with some domains limited to January 2025, according to the DeepMind model card.

So, is Gemini 3.7 Flash worth it?

Gemini 3.7 Flash is a real upgrade for agentic tasks: AutomationBench nearly doubled, FrontierCode went up nine points, execution got more disciplined. That's true.

It's also true that it didn't get cheaper, hallucinates more than its predecessor, loses to GPT-5.6 Terra on the overall index, and doubles in price on a date that's already set.

Both things hold at once. That's why the answer to "is it worth it?" doesn't come from the release or from this post: it comes from your eval set running on your domain. A new model every week is the new normal. Having a way to measure is what separates those who decide from those who react.

Lucas Souza
Written by
Lucas Souza

{AI Engineer} — apaixonado por Laravel, arquitetura de software e construir produtos com impacto. Compartilho aqui tutoriais, descobertas e reflexões sobre o dia a dia de engenharia.

▪ Clã Beer and Code

There is no shortage of content. What is missing is someone to untangle it: what matters now is how to implement it the right way. In the Clã you get that live, every week, with people who have already filtered out the noise.

Join the Clã
Meet the Clã Beer and Code
playing