Gemini is the punchline of AI and the cheapest way to run an agent
Google's 3.7 Flash matches Claude and GPT on composite intelligence at roughly a third of the price — but the number holding that argument up expires on 31 December, and the footnote saying so is on Google's own table.
Gemini has been the internet's favourite AI joke for about three years. It counted the letters in "strawberry" wrong in public. Its image model was withdrawn and apologised for. Every launch since has arrived to a chorus of screenshots of the one thing it got stupidly, obviously wrong, and the jokes have been good, because the failures were real.
Gemini 3.7 Flash, which Google shipped on 13 August, will not end that. It came out and lost the benchmark everyone quotes. But the numbers Google published alongside it are the clearest statement yet of what the company has actually been doing while being laughed at, and it is not trying to win the argument the jokes are about.

The chart that is the actual announcement
Google published six charts. Five are bar charts of the sort every lab ships. The sixth is a scatter plot, it is not Google's data, and it is the only one that says anything a competitor cannot answer with a bar of their own.

It plots score on DeepSWE v1.1 — long-horizon software engineering, the closest thing the industry has to a test of whether a model can close a real ticket — against what one task costs to run. The frontier sweeps from Claude Opus 5, top left, buying about 73 percent at over ten dollars a task, down to the cheap models at the right that cost almost nothing and deliver accordingly.
Three dots are highlighted in blue: gemini-3.5-flash, gemini-3.6-flash, gemini-3.7-flash. They form a trajectory, and the direction is the point. Each generation moves up and to the right — better *and* cheaper — until 3.7 Flash lands on the efficient frontier itself, at roughly 65 percent for around four dollars a task, within reach of models costing three times as much.
That is the announcement, and it is not that Google has built the smartest model. It is that Google has moved the price of a competent agent again, on schedule, for the third generation running.
The benchmark that gets screenshotted
Here is the same benchmark drawn the way it circulated on X for the last four days.

GPT-5.6 Terra: 69.6 percent. Gemini 3.7 Flash: 65.3. A 4.3-point loss on the headline agentic coding test, from the company that has been promising to catch up for two years. The screenshots wrote themselves, and they were not unfair — Google chose to publish this chart, and on it, Google loses.
What the bar chart removes is the x-axis. The scatter plot and the bar chart contain the same DeepSWE result; one of them mentions that the winning model costs several times more per task, and the other is the one that travelled.
This is not a defence of Gemini so much as an observation about how model news moves. A bar chart compares one dimension and fits in a screenshot. A scatter plot compares two and requires a caption. The industry has standardised on the format that hides the variable most enterprises actually optimise.
The chart nobody screenshotted
If you want to know what Google thinks it is selling, skip the coding benchmarks and look at the one with the ugliest name.

AutomationBench measures enterprise workflow automation — the unglamorous business of moving a task through six systems without a human. Gemini 3.7 Flash scores 30.4 percent. GPT-5.6 Terra scores 23.6. Claude Sonnet 5 scores 10.7.
That last number deserves a second look. Sonnet 5 is a genuinely strong model — it beats Gemini on Agent's Last Exam and on the human-solvable half of BioMysteryBench — and on this benchmark it scores roughly a third of what Gemini does. Whatever AutomationBench is measuring, Anthropic is not optimising for it and Google is optimising for almost nothing else.
The generational jump says the same thing louder. Gemini 3.6 Flash scored 17.0 on AutomationBench six months ago. 3.7 Flash scores 30.4 — a 79 percent improvement, at the same list price. The interesting comparison here was never Gemini against Claude. It is Gemini against Gemini two releases ago.
The footnote that eats the argument
Now the part that should change how you read every chart above.

The prices on Google's table — $0.75 per million input tokens, $3.75 per million output — carry an asterisk. The note reads: introductory pricing expires on 31 December 2026. From 1 January, the rate is $1.50 and $7.50. Double, on both axes.
Run that through the comparison. Today, against Claude Sonnet 5 at $2.00 and $10.00, Gemini 3.7 Flash is about 2.7 times cheaper on input and output alike. In January it is 1.3 times cheaper. The gap does not close because a competitor cut prices; it closes because the introductory period ends.
Against Muse Spark 1.2 it does not merely close. Muse Spark lists at $1.25 and $4.25 — after 1 January, Gemini 3.7 Flash is *more expensive* on both input and output than a model that scores one point higher on the Artificial Analysis intelligence index. Every cost-efficiency claim in this article, including the scatter plot that anchors it, has a four-and-a-half-month shelf life.
None of which is hidden. It is printed on Google's own comparison table, in grey, under an asterisk, and it has been almost entirely absent from four days of coverage arguing about a 4.3-point gap on DeepSWE.
What the demos are arguing
Google shipped video with the launch, and the demos are worth watching for what they assume rather than what they show.
A PDF becoming an interactive data story. A landing page generated in one shot by a model orchestrating sub-agents. A 3D game with characters and textures generated in real time. A robotics model trained inside a three-agent graph loop.
Not one of them is a single clever answer. Every one is a loop — a model called repeatedly, in a graph, generating and revising until something works. That is a workload where price per token is not a line item but the entire constraint, because the token count is unbounded by design. A model twice as good and three times the price loses that argument on arithmetic.
Which is presumably why Google put the scatter plot in the post at all.
Where it is honestly worse
An analysis that only found supporting evidence would not be worth reading, so: Gemini 3.7 Flash loses, clearly, in several places on Google's own table.
It loses long-horizon agentic coding to GPT-5.6 Terra on DeepSWE (65.3 to 69.6), on Terminal-bench 2.1 (85.8 to 87.4) and badly on Terminal-bench 3.0 (14.9 to 20.8). It loses knowledge work to Muse Spark 1.2 on GDPVal-AA by 103 Elo — 1525 against 1628 — which is not a rounding error. It loses Agent's Last Exam to Claude Sonnet 5, 26.3 to 33.3, and the hard half of BioMysteryBench to Terra, 43.5 to 49.4.
And on CharXiv Reasoning with tools it scores 88.7 against its own predecessor's 89.4. A regression, published without comment, on the company's own comparison table.
The shape of those losses is consistent. Where the task is long, hard and unusual — the tasks you would pay a premium for anyway — Gemini is second or third. Where the task is repetitive, structured and high-volume, it leads, sometimes by a lot.

Look at the absolute numbers on that one before drawing any conclusion about who is winning. The best score on expert PDF comprehension, across every model tested, is 34 percent. Two-thirds of the time the leading model reads a hard document wrong. On Terminal-bench 3.0 the field's best is 20.8 percent. The gaps everyone is arguing about sit on top of failure rates that would end a pilot in any other industry.
What to do with this

Three things follow, and none of them is "Gemini is back".
If you are running high-volume agentic work — document extraction, workflow automation, code review at scale — 3.7 Flash is the strongest price-performance argument on the table today, and you should be measuring it against your own workload rather than against a bar chart. If your work is long-horizon and genuinely hard, Terra and Opus still win the tasks and the scatter plot tells you what that costs.
And whatever you conclude, put 1 January 2027 in the calendar. The most persuasive number in this launch is scheduled to double, and a procurement decision made on the introductory rate is a decision made on a price that has an expiry date printed next to it.
The jokes about Gemini will continue, and some of them will keep being deserved. But the company has now shipped three consecutive generations that move the same dot up and to the right on the same chart, and it published that chart rather than the one where it wins. Being laughed at while systematically collapsing the cost of the thing everyone else is selling at a premium is a strange position to be in. It is not obviously a losing one.
Sources
This article was written from these pages. Read them.
- primaryIntroducing Gemini 3.7 Flashblog.google
- primaryGemini 3.7 Flash evaluation methodologydeepmind.google
- primaryDeepSWE v1.1 leaderboard and cost-per-task datadeepswe.datacurve.ai
Analysis written and edited by a person at Epoch, from the primary announcement and the published benchmark figures. Charts are the subject's own; the reading of them is ours.