What a 1 Million Token Output Limit Means
A 1 million token output limit is not a bigger context window. Here is what the output budget actually caps, why it is lower, and what it costs.
A 1 million token output limit is the cap on how much text a model can produce in one response. It is not the context window, which caps how much it can read. Those two numbers are usually different by an order of magnitude or more, they are the limits people confuse most often, and only one of them is the reason your long generation stopped halfway through a sentence.
Google's Gemini 4 Argon put the figure in the news by raising its output ceiling from 64,000 tokens to 1,000,000. The concept is older than the headline and applies to every model you use.
Two budgets, not one
Every API call has two separate allowances:
Context window | Output limit | |
|---|---|---|
What it measures | Everything the model reads: system prompt, history, documents, plus what it writes | Only what the model writes in this response |
Typical size in 2026 | 200K to several million tokens | 8K to 128K on most models |
What happens at the limit | Oldest content is dropped or the call is rejected | Generation stops mid-sentence |
Who usually hits it | Long chats, big document loads | Anyone asking for a long artifact in one pass |
The output limit is almost always the smaller of the two, and it is nested inside the first. If a model has a 1M context window and a 64K output limit, you cannot spend 900K on input and still get 64K out unless the window has room for both.
If you are here because a response stopped mid-sentence and you want the diagnosis rather than the concept, read why AI responses get cut off instead. This piece is about what the number on the spec sheet buys you.
What a 1 million token output limit buys in practice
Tokens are not words. For ordinary English prose, a useful conversion is about 0.75 words per token, so:
64,000 tokens is roughly 48,000 words, about a short novel
1,000,000 tokens is roughly 750,000 words, about seven long novels
Code runs denser, closer to 0.4 to 0.5 words per token, so the same budget yields fewer lines than prose
Stated that way, the old 64K ceiling sounds generous. It is not, because the tasks that hit it are not essays. They are structured artifacts: a refactor that touches nine files, a migration plus its rollback plus the test file, a translated document that must come back whole. Each of those has a minimum size set by the work, not by how verbose the model feels.
Why output limits are so much lower than context windows
Reading and writing cost differently. Input tokens can be processed in parallel across the whole sequence. Output tokens are produced one at a time, each conditioned on everything generated before it, so the work is sequential and cannot be batched away. That is why output tokens cost more than input tokens on essentially every provider's price list, often by a factor of four or five.
It is also why the limits exist as product decisions rather than physics. A single request that generates a million tokens occupies a slot for a very long time. Providers cap output partly to keep latency predictable for everybody else.
The arithmetic you should run before relying on a big output budget
Three numbers decide whether a large output limit is useful to you or just a line on a spec sheet.
Cost per maximum response. At Argon's introductory $10 per million output tokens, one full-length response is about $10. Fine once a day, ruinous in a retry loop.
Time to finish. At a generous 100 tokens per second, a million tokens takes around three hours of continuous streaming. Your HTTP client, your load balancer and your user all have opinions about that.
Coherence over the distance. No public benchmark currently measures whether a model stays consistent across 500,000 tokens of its own output. Until one does, treat long-generation quality as an open question rather than an advertised feature.
A quick sanity check in code, using the rough prose ratio:
WORDS_PER_TOKEN = 0.75
PRICE_PER_M_OUT = 10.00 # dollars, Argon introductory rate
TOKENS_PER_SEC = 100 # optimistic
def budget(tokens):
words = int(tokens * WORDS_PER_TOKEN)
cost = tokens / 1_000_000 * PRICE_PER_M_OUT
minutes = tokens / TOKENS_PER_SEC / 60
print(f"{tokens:>9,} tokens ~{words:>7,} words ${cost:>6.2f} {minutes:>6.1f} min")
for t in (8_000, 64_000, 128_000, 1_000_000):
budget(t)The 8,000 token row is the one most teams are actually living in, and it is where the honest answer is usually to split the task rather than to buy a bigger ceiling.
When a bigger output limit genuinely helps
Artifacts that cannot be split without losing consistency: a schema migration and its rollback, a document translation where terminology must match throughout.
Agentic runs where the model writes its reasoning as it works and the trace counts against the output budget.
Batch jobs with no human waiting, where three hours of streaming is acceptable.
And when it does not: anything interactive, anything retried, and anything you could decompose into independent chunks and assemble yourself. Chunking is cheaper, faster, easier to debug, and each piece fits comfortably in an 8K budget.
Common questions
Is the output limit included in the context window?
Usually yes. Most providers count generated tokens against the same window as the input, so a long response eats room you might have wanted for context. Check the specific model's documentation rather than assuming, because the accounting differs.
How do I know which limit I hit?
The API tells you, via the finish reason on the response. That diagnosis has its own guide: mapping each truncation symptom to the field that proves its cause is the troubleshooting counterpart to this piece.
Does a higher output limit mean a smarter model?
No. It is a capacity parameter, like disk size. It changes what you can ask for in one call, not the quality of what comes back. The same question applies here as to context windows, which is covered in does a bigger context window mean better answers.
Can I just set max_tokens to the maximum and forget about it?
You can, and you probably should not. max_tokens is your cost ceiling as well as your length ceiling. Setting it to a million means a single malformed prompt that triggers a runaway generation bills you for a million tokens.
Where does this sit relative to tokens generally?
If the word token is still doing heavy lifting in your head, start with what is a token in AI and then what is a context window. The release that prompted this piece is covered in our write-up of Gemini 4 Argon, and Google's announcement is the primary source for the figures.
How did this land?
About the author

Senior Editor, AI & Product
Cecilia leads the Swarmz editorial desk. She has spent a decade turning complex AI and product topics into writing people actually finish, and she owns the blog's quality bar.


