Dashboard

Gemini 4 Argon: What Google Actually Shipped

Google released Gemini 4 Argon on 30 September with a 1M token output limit. Here is what changed, what the benchmarks say, and who can use it.

Cecilia Iona
Cecilia Iona
Senior Editor, AI & Product
1 October 20261 min read

Google released Gemini 4 Argon on 30 September 2026, and the number worth writing down is not a benchmark score. It is 1,000,000. That is Argon's output token limit, up from 64,000 on the previous generation, a roughly 15x jump in how much a single model call can write before it stops.

The second thing worth writing down is that you probably cannot use it yet. Gemini 4 Argon is rolling out first to a closed set of cyber defenders, not to the API.

What Google announced

Per Google's own announcement, Argon is built for what it calls long-horizon workflows: software engineering, legal and finance knowledge work, and defensive cybersecurity. Google says it can locate a software vulnerability, validate that it is real, and write the patch without a human in the loop. The demonstration it leads with is a critical flaw exposing personal data in healthcare software used by hospitals, found by the security firm Wiz using Argon, which Google says earlier frontier models had missed.

The published figures, from Google's post:

Benchmark

Argon result

DeepSWE v1.1

77.9%

CWE-bench v1

68%, tied for first

LVBench (long video)

91.7%

AutomationBench

51.3%, ranked first

Output token limit

1M, up from 64K

Introductory price

$2 per million input, $10 per million output

Google also claims a lead on the Gray Swan prompt injection robustness benchmark without publishing a number for it. Treat the unnumbered claim differently from the numbered ones.

Who can actually use Gemini 4 Argon

Access is staged, and the first stage is narrow. Argon goes first to trusted defenders through Google's Fairwind Program, described as controlled access for governments and critical-infrastructure operators, with terms restricting use to defensive and research work. Help Net Security reports developers, enterprises and consumers come later, starting with paid API customers and Google AI Ultra subscribers. No date has been attached to that second stage.

So the practical answer for anyone building on top of a model today is: nothing changes this week. You cannot route traffic to Argon, you cannot price against it, and you cannot benchmark it on your own work.

Why the output limit is the interesting part

Context windows have been growing for two years and the industry has largely stopped being surprised by them. Output budgets have not moved nearly as fast, and they are the limit people actually hit. A 64,000 token ceiling is roughly 45,000 words of generated text, which sounds enormous until you ask a model to refactor a module, write the migration, and produce the test file in one pass.

Raising that ceiling to a million tokens changes the shape of the task you can hand over, not the quality of the answer. It is the difference between asking for a plan and asking for the finished work. Whether a model can stay coherent across 1M tokens of its own output is a separate question, and one no public benchmark answers yet.

At $10 per million output tokens, a single maximum-length response costs about $10. That is cheap for a day of engineering and expensive for something you call in a loop, which is the tension every team adopting this will have to price.

What to watch next

  • Whether the API tier arrives with the same 1M output ceiling or a reduced one. Staged launches often ship the headline number only to the top tier.

  • Whether independent evaluators reproduce the CWE-bench and DeepSWE figures. Google cites the benchmarking firm Vals, and TechCrunch's write-up notes the comparison set includes OpenAI's GPT-6 Astra and recent Anthropic models.

  • Whether the Fairwind restrictions hold. A model Google says can autonomously write patches can autonomously write the other thing, which is presumably why the rollout looks like this.

If you track releases to decide when to switch models, this one is not actionable yet, and that is a useful thing to know early. Our guide to reading an AI release note without the hype covers the questions that separate a shipped capability from an announced one, and telling whether a model release is actually a big deal covers the follow-up. For the broader habit of keeping up without drowning, start with how to keep up with AI news.

The context window side of the same story is covered in what a 4 million token context window means, and the cost asymmetry Argon's pricing reflects is explained in why output tokens cost more than input tokens.

How did this land?

About the author

Cecilia Iona
Cecilia Iona

Senior Editor, AI & Product

Cecilia leads the Swarmz editorial desk. She has spent a decade turning complex AI and product topics into writing people actually finish, and she owns the blog's quality bar.

Share

Get the next post in your inbox

One email a month. Product updates, engineering posts, and the best of Built with Swarmz.

I agree to receive emails about AI building tips and Swarmz product news. Unsubscribe any time.