Gemini 3.7 Flash went live today, August 13, 2026, and the pitch is blunt: better than 3.6 Flash almost everywhere, at half the launch price. That combination is rare. Model releases usually trade one for the other. Below are the five things that actually matter if you build with these models: the coding gains, the pricing (including the January 2027 catch), agent performance, hard specs, and where you can run it today. Every figure comes from Google’s own model card and launch documentation.
Here’s the fast comparison before we get into it:
| Metric | Gemini 3.7 Flash | Gemini 3.6 Flash | Claude Sonnet 5 | GPT-5.6 Terra |
|---|---|---|---|---|
| Input price ($/1M tokens) | $0.75* | $0.75* | $2.00 | $2.00 |
| Output price ($/1M tokens) | $3.75* | $3.75* | $10.00 | $12.00 |
| FrontierCode 1.1 (production code) | 43.6% | 34.4% | 42.7% | 41.3% |
| DeepSWE v1.1 (long-horizon SWE) | 65.3% | 48.6% | 53.8% | 69.6% |
| Terminal-bench 2.1 (agentic coding) | 85.8% | 78.0% | 80.4% | 87.4% |
| AutomationBench (enterprise workflows) | 30.4% | 17.0% | 10.7% | 23.6% |
| WebDev Arena (Elo) | 1588 | 1538 | 1541 | 1523 |
Introductory pricing through December 31, 2026. It doubles on January 1, 2027.
1. The coding gains are the headline, and they hold up
The biggest single jump is in software engineering. On FrontierCode 1.1, which measures production code quality, 3.7 Flash scores 43.6% against 34.4% for 3.6 Flash. On DeepSWE v1.1, the long-horizon software engineering benchmark, it hits 65.3% versus 48.6%. That is a 16-point jump in three weeks, since 3.6 Flash only shipped on July 21, 2026.
“Our most intelligent workhorse model yet for coding and agents,” is how Tulsee Doshi, Senior Director of Product Management, framed it in Google’s launch announcement.
The practical claims behind those numbers: higher first-pass code accuracy, more functional web layouts in fewer prompts, and closer adherence to UI design specs when you feed it a screenshot or a full design system. Its WebDev Arena Elo of 1588 beats 3.6 Flash by 50 points.
Now the honest part. It is not the best coding model on the market. GPT-5.6 Terra still leads it on DeepSWE (69.6%) and Terminal-bench 2.1 (87.4%). What 3.7 Flash does is beat Claude Sonnet 5 on FrontierCode, AutomationBench, and WebDev Arena while costing roughly a third as much per output token. For teams running thousands of agent turns a day, that ratio is the story, not any single benchmark.
Curious what AI could do for your business?
No jargon and no hard sell. Just a friendly look at where AI fits, and where it doesn't.
2. What does Gemini 3.7 Flash actually cost?
Gemini 3.7 Flash costs $0.75 per million input tokens and $3.75 per million output tokens at launch, with those introductory rates doubling to $1.50 and $7.50 on January 1, 2027, per Google’s Gemini API pricing page.
That’s half of what 3.6 Flash launched at three weeks ago. And here’s what nobody puts in the headline: thinking tokens bill as output tokens. An agent that reasons hard on a “high” thinking budget can burn far more output than the visible response suggests, so model your costs on real traces, not on response length.
The rest of the price sheet, as of August 13, 2026:
- Context caching: $0.075 per million tokens during the introductory window, plus $0.50 per million tokens per hour of storage. Both double in 2027.
- Batch and Flex tiers: 50% off standard input and output rates.
- Grounding with Google Search and Maps: 5,000 free requests per month shared across Gemini 3.x models, then $14 per 1,000 requests. That fee stings at scale if grounding is on by default.
One planning note from teams that run inference budgets for a living: build your unit economics on the January 2027 prices. If your margin only works at $0.75 input, it doesn’t work.
3. It’s aimed squarely at agents and enterprise workflows
Skip the coding numbers for a second. The eval where 3.7 Flash separates most from the field is AutomationBench, a private-set benchmark for enterprise workflow automation: 30.4% against 17.0% for 3.6 Flash, 23.6% for GPT-5.6 Terra, and 10.7% for Claude Sonnet 5. Google’s model card also reports 90.7% on Harvey LAB-AA for complex legal workflows, 34.0% on GDP.pdf for expert document comprehension (up from 22.0%), and 97.0% on GDM-MRCR v2 long-context retrieval at 128k.
The qualitative claim matches what agent builders care about: more disciplined multi-step execution, fewer retries when the agent hits an obstacle, and better instruction fidelity. Less babysitting, in plain terms. Named early testers include Harvey, Hebbia, LangChain, Cartwheel, Nunu.ai, Open Code, and Stanford’s Department of Biology.
It’s not a clean sweep, though. On GDPVal-AA v2, the general knowledge-work Elo, 3.7 Flash posts 1525 while Claude Sonnet 5 sits at 1598 and GPT-5.6 Terra at 1578. So for open-ended analyst work, the pricier models still hold an edge. For structured, repeatable business workflows, this model’s cost-per-completed-task math is hard to argue with.
4. Specs and limits to know before you build
The core numbers, from Google’s developer documentation, model code gemini-3.7-flash:
- Context window: 1,048,576 input tokens.
- Output limit: 65,536 tokens.
- Inputs: text, image, video, audio, and PDF. Outputs: text only. No native image or audio generation.
- Tooling: function calling, structured outputs, code execution, computer use (preview), file search, URL context, caching, Batch API, and Flex and Priority inference tiers.
- Thinking budget: adjustable low, medium, or high, so you pick the quality-cost-latency mix per request.
The knowledge cutoff deserves a careful read. It’s March 2026 for most domains, but some areas remain limited to January 2025, the same split structure 3.6 Flash used. If your workflow touches fast-moving facts, wire in grounding or retrieval rather than trusting the weights. Google’s own model card also flags hallucinations and occasional slowness or timeout issues as known limitations, which matters if you’re putting this behind a latency SLO. Plan retries.
5. Where to get it, and what the safety review found
This is a general availability release, not a preview. As of today it’s live in Google AI Studio, the Gemini API, Android Studio, the Gemini Enterprise Agent Platform, Google Antigravity (Google’s agent-first coding IDE), and Gemini Spark for AI Pro and Ultra subscribers in more than 160 countries. It lands two days after the Gemini app passed 1 billion monthly active users on August 11, 2026.
On safety: Google evaluated the model against the April 2026 version of its Frontier Safety Framework across CBRN, cybersecurity, harmful manipulation, and ML R&D and misalignment. It reached no tracked or critical capability levels. Worth knowing: it did hit the alert threshold, one step below the critical line, in both CBRN and cybersecurity, and Google says it’s shipping updated safeguards in both domains. Child-safety evaluations met required launch thresholds, and external red teaming found no egregious concerns.
Zoom out and the cadence is the real signal. Four Flash releases in under eight months: Gemini 3 Flash on December 17, 2025, then 3.5 Flash in May 2026, 3.6 Flash on July 21, and now 3.7. Whatever model you standardize on this quarter will be superseded next quarter. Architect for that.
FAQ
How is Gemini 3.7 Flash different from Gemini 3.6 Flash?
Same architecture, better reasoning. Google describes it as algorithmic improvements to the core reasoning foundation of 3.6 Flash rather than a new architecture, with gains concentrated in code generation, web development, document-heavy knowledge work, and multi-step agentic execution. It launched at half of 3.6 Flash’s original per-token price.
What is the context window of Gemini 3.7 Flash?
1,048,576 input tokens, roughly 1 million, with a 65,536-token output limit. On the GDM-MRCR v2 long-context benchmark it scores 97.0% at 128k, up from 91.8% for 3.6 Flash.
Can Gemini 3.7 Flash generate images or audio?
No. It accepts text, images, video, audio, and PDFs as input, but outputs text only. Google’s launch demos pair it with other models, such as Nano Banana for image generation, when a workflow needs visual output.
What is the knowledge cutoff for Gemini 3.7 Flash?
March 2026 for most domains. Some domains are limited to January 2025, consistent with the rest of the Gemini 3 family, so grounding or retrieval is the safer bet for current-events work.

What could a custom AI agent take off your plate?
We build production-grade AI systems that quietly handle the busywork, so your team can focus on the work that actually matters.
How to evaluate Gemini 3.7 Flash for your own stack
Don’t take the benchmark table’s word for it. Pull 50 to 100 real traces from your current agent or pipeline, replay them against gemini-3.7-flash at each thinking budget, and compare completion rates and total token spend, thinking tokens included. Then rerun the math at the January 2027 prices before you commit. The teams that get burned are the ones who standardize on introductory pricing and demo-grade evals.
If you’d rather not run that gauntlet alone, this is exactly the work we do at AlphaCorp AI, from production AI agent builds to model-swap evaluations on live workloads. Talk to the people who’d actually build it.





