Skip to content
Carlos KiK
Go back

Gemini 3.7 Flash Is The Workhorse Bet

Three weeks. That is how long Gemini 3.6 Flash lasted before Google replaced it.

Gemini 3.7 Flash is here, and Google is calling it its most intelligent workhorse model yet for coding and agents. The benchmark gains are real: 43.6% on FrontierCode 1.1 Main versus 34.4% for 3.6, 65.3% on DeepSWE v1.1 versus 49.0%, and a jump on WebDev Arena from an Elo of 1538 to 1588.

But the number that matters is the price.

Google is offering 3.7 Flash at an introductory $0.75 per million input tokens and $3.75 per million output tokens through the end of the year, half the original cost of 3.6 Flash. That is not a discount. That is a statement about where the volume of agent work is going.

The workhorse is the product

Frontier models do the heroics. They get the demos, the press releases, the “look what it proved” moments. A production system does not run on heroics, though. It runs on the model that handles the ten thousand routine steps between the interesting ones, the edits, the extractions, the tool calls, the retries, the boring multi-step work that has to be right more often than it has to be brilliant.

That is the job 3.7 Flash is built for, and the pricing says Google knows it.

The developer experience notes are worth reading closely. Google says 3.7 Flash adapts to roadblocks, clarifies intent when needed, and puts in more effort on multi-step planning and tool calls. In plain terms, it is less likely to wander and less likely to need a human to catch it. For an agent that runs unattended for hours, that difference is the difference between a tool and a liability.

The quiet part is Gemini Spark. Google’s 24/7 personal agent for AI Pro and Ultra subscribers is moving to 3.7 Flash starting today, in over 160 countries. The personal agent layer is being routed to the cheap workhorse, which tells you what Google thinks the marginal agent task actually costs.

I wrote about Gemini 3.6 Flash landing in Copilot as a model routing signal. This is the same signal from the other side: the model being routed. When the workhorse gets twice as capable at half the price, the routing decision gets easier, and the interesting question stops being which model is the smartest.

It becomes which workflow you own around it.

Source: Google


Share this post on:

Previous Post
Relay Is Shutting Down, and Its Founder Is Joining Chrome
Next Post
Muse Glimmer Brings The Model Back To The Personal Computer