Gemini 3.8 Flash: The $0.75 Model That Beats Opus 5 (On Some Rows)
Gemini 3.8 Flash: The $0.75 Model That Beats Opus 5 (On Some Rows)
September 2, 2026 · by OpenSource Factory · pricing and benchmarks verified against Google's announcement on the day of release
Three weeks ago I wrote about Gemini 3.7 Flash — half the price, still no 3.5 Pro. Today Google did it again: Gemini 3.8 Flash is out, plus a "Flash Cyber" variant, three weeks after 3.7 and just six weeks after 3.6. That's three Flash releases in six weeks. Google has stopped treating Flash like a model and started treating it like a service that gets better every time you open it.
Gemini 3.8 Flash and 3.8 Flash Cyber — the reasoning-and-coding push, three weeks after 3.7. (image: Google)
The headline: a $0.75 model beating $5 and $25 models
Intro pricing holds at $0.75 input / $3.75 output per million tokens through December 31 (then $1.50/$7.50, same pattern as 3.7). Claude Opus 5 costs $5/$25. GPT-5.6 Sol costs $4/$20. So the Flash is six to seven times cheaper than the flagship it's now beating on real workloads — and Google's own published table shows wins on finance, legal, video, biology, and chart reasoning.
Where the $0.75 model wins. Vals Finance 61.4 vs Opus 5's 58.6; Harvey Legal 10.0 vs GPT-5.6 Sol's 2.5 (four times); LVBench 87.8 vs Opus 5's 75.4. Vendor-reported, day-1, redrawn.
The legal number deserves a second look. Harvey's Legal Agent Benchmark — complex legal workflows — shows Flash at a 10.0% pass rate versus GPT-5.6 Sol's 2.5% and Opus 5's 6.7%. Four times the OpenAI flagship at a sixth of the price. Finance (Vals Agent v2) and long-video understanding (LVBench, 87.8 vs 75.4) are similarly decisive. This is Google going after the workflows enterprises actually bill for — not just the coding leaderboards.
The honest gap: where the cheap Flash still loses
Google isn't claiming an outright win, and neither should we. The flagship gap shows up exactly where you'd expect — general agentic ability and long-horizon work. The standout is Terminal-bench 4.0: 19.1% for Flash vs 51.8% for Opus 5. That's not a close race; that's a different class of general-purpose agent. OSWorld computer use (59.0 vs 75.4), knowledge-work Elo (1545 vs 1824 on GDPval), and DeepSWE long-horizon software engineering (71.0 vs 74.0) all still favor Opus 5.
The honest read: on general agentic ability, the flagship is still the flagship. Terminal-bench 4.0's 51.8 vs 19.1 is where a cheap fast model and a frontier reasoning model diverge most. Vendor-reported, redrawn.
Against its own predecessor, though, 3.8 Flash is better almost everywhere: Terminal-bench 4.0 up from 11.2, OSWorld up from 50.6, BioMystery's hard tier up from 43.5 to 56.5. The model card describes it as built on 3.7 Flash with gains in software engineering and agentic knowledge workflows — an iteration, not a reinvention. Same 1M context, same 64K output, effort levels intact.
Google's own eval table — price on top, wins shaded. Flash leads most rows but Opus 5 takes GDPval, Terminal-bench 4.0, and OSWorld. (image: Google, via 9to5Google)
The price story is the strategy
Look at the chart below and the pattern is unmistakable. Google isn't trying to out-frontier Anthropic at the top of the market. It's pricing Flash so far below the flagships — while getting within striking distance on the workflows that matter — that for any business running agents at scale, the decision stops being about benchmarks and becomes arithmetic.
Six to seven times cheaper per token than Opus 5. Intro pricing holds to Dec 31 2026, then $1.50/$7.50. Vendor list prices, redrawn.
My take
The Flash cadence — three releases in six weeks — is the real story, and it reframes what I said about 3.7. Back then the angle was "half the price of the frontier, still no flagship." Now the angle is sharper: the half-price model is beating the flagship on finance, legal, video, and biology rows, at a sixth of the cost. If you're running high-volume agentic workloads in those domains, the case for Opus 5 gets hard to defend. If your work is general-purpose agentic coding where the last points decide completion — Terminal-bench 4.0 territory — the flagship gap is still very real, and Opus 5 remains the pick.
One caution, same as always: these are day-1 vendor numbers on Google's own table, and the same-day release pattern means rival models' scores may predate this snapshot. The "Flash Cyber" variant (safeguards-tuned, like the Mythos pattern across the industry) tells you where the frontier is heading. I'll update when independent leaderboards land. Also worth noting for the pricing calendar: the intro rate expires December 31, and if history is any guide, the $1.50/$7.50 revert will be the follow-up story — the same one I flagged for 3.7.
Sources: Google announcement (X) · DeepMind model card · 9to5Google · officechai · benchmarks vendor-reported as of September 2, 2026. Related: my Gemini 3.7 Flash post.
Comments
Post a Comment