Gemini 3.8 Flash: Third Flash Model in Six Weeks
Google shipped Gemini 3.8 Flash on 2 September, three weeks after 3.7. The intro price holds at $0.75 per million input tokens until 31 December, then doubles.
Google shipped Gemini 3.8 Flash and 3.8 Flash Cyber on 2 September 2026, three weeks after 3.7 Flash and the third Flash release in six weeks. Pricing holds at the same introductory $0.75 per million input tokens and $3.75 per million output tokens. The detail worth writing down is in the pricing footnote: after 31 December 2026 those rates become $1.50 and $7.50, a straight doubling with a date attached.
Put 31 December in your cost model now
Introductory pricing with a published expiry is unusual enough to be useful. Most providers change prices when they feel like it, and you find out from a billing alert. Here you have four months of notice and an exact figure, which means there is no excuse for the January surprise.
If Gemini 3.8 Flash is going into a product, run two cost lines from the start: one at the intro rate and one at the post-December rate. If the second line does not work, you are building on a price that expires, and the time to find that out is now rather than in a January board meeting. We cover the general version of this problem in what to do when your AI provider raises prices.
What the model is actually better at
Google frames 3.8 Flash as its most intelligent workhorse model, with gains concentrated in software engineering, agentic tasks and multi-step reasoning in specialist domains. The published figures:
DeepSWE v1.1, a long-horizon software engineering benchmark, where Google says 3.8 Flash beats most larger frontier models at a fraction of their cost.
HLE-Verified at 54.9%, covering multi-step reasoning across STEM, humanities and professional fields.
Vals Finance Agent V2 and Harvey's Legal Agent Benchmark, both domain-specific agent evaluations rather than general knowledge tests.
Gray Swan, which measures robustness against prompt injection, described as a significant leap.
The Gray Swan line deserves more attention than it will get. Prompt injection robustness is the thing standing between an agent with tool access and an expensive incident, and it is rarely the number anyone quotes from a launch post. If you are letting a model touch real systems, that metric matters more to you than the reasoning score does.
Availability is broad: Google AI Studio, Android Studio, Gemini Enterprise, plus the Gemini app, Search AI Mode and Google Sheets for AI Pro and Ultra subscribers.
Flash Cyber and the Fairwind Program
The Cyber variant is not on general release. It goes to what Google calls trusted defenders through the Fairwind Program: government authorities, critical infrastructure operators and software maintainers. The numbers Google published are the most concrete vendor claims we have seen for security work:
Claim | Figure | Source |
|---|---|---|
CWE-Bench patching | 47.2% pass@1 | |
Internal multilingual benchmark | Over 70% success across 20 languages | |
Correct patches vs leading commercial models | 2.6 times more | Chrome Security team |
Vulnerability recall | 7.5% to 9.7% higher at 2.3x to 5.2x lower cost | Wiz |
Critical vulnerability discovery | Under 2 hours | Google Cloud Vulnerability Research |
Third-party attribution on two of those rows is what makes them interesting. A vendor quoting its own benchmark is a marketing claim; a vendor quoting the Chrome Security team and Wiz is at least a claim someone outside the model team put their name to. We look at what these numbers mean for ordinary teams in can AI find security vulnerabilities in your code.
The cadence problem
Three Flash releases in six weeks is the real story for anyone maintaining production code. Every release resets your evaluation work, and the honest answer for most teams is that you cannot re-qualify a model every 21 days. Pin a version, decide in advance what would make you move, and check on a schedule you choose rather than the one Google publishes on.
The previous release in this sequence is covered in our write-up of Gemini 3.7 Flash, which is worth reading side by side to see how little the price moved and how much the benchmark framing did. On the wider question of how often it is rational to switch, see how often should you switch AI models and, if the cost line is what is driving you, how to reduce AI API costs. The general habit of not chasing every launch is in how to keep up with AI news.
How did this land?
About the author

Senior Editor, AI & Product
Cecilia leads the Swarmz editorial desk. She has spent a decade turning complex AI and product topics into writing people actually finish, and she owns the blog's quality bar.


