Dashboard

Why Output Tokens Cost More Than Input Tokens

Output tokens cost three to six times more than input tokens. The reason is prefill versus decode, and it changes which half of your prompt to cut.

Cecilia Iona
Cecilia Iona
Senior Editor, AI & Product
12 September 20261 min read

Why Output Tokens Cost More Than Input Tokens

Output tokens cost more than input tokens because of how the hardware runs, not how the vendor prices. Your whole prompt goes through the model in one parallel pass. Every token of the answer needs its own pass, one after another, each waiting for the one before it. Reading is a batch job. Writing is a queue.

On current price sheets that difference typically lands somewhere between three and six times.

Prefill reads in parallel, decode writes in single file

A model call has two phases with genuinely different shapes.

Prefill processes your prompt. Every token is available at once, so the GPU does one large matrix operation across the whole sequence. Two thousand input tokens do not take two thousand times as long as one. They take a bit longer than one thousand, and the hardware stays busy the entire time.

Decode produces the answer. The model generates one token, appends it to the sequence, and runs again to work out the next. Token 400 cannot be computed before token 399 exists. Each step is a full pass through the model weights to produce a single token, which is a terrible ratio of work moved to work done. The GPU spends most of decode waiting on memory rather than computing.

That is the whole explanation. Output tokens occupy the expensive, badly parallelised half of the call.

The ratio, on real price sheets

Model

Input per 1M

Output per 1M

Ratio

Sakana Fugu Max

$2

$6

3x

Sakana Fugu Ultra v2

$5

$30

6x

DeepSeek V4.1-Flash (off-peak, cache miss)

RMB 1

RMB 4

4x

Figures from the Sakana Fugu release and the DeepSeek V4.1-Flash announcement, both published this week. Cached input is cheaper again, often by an order of magnitude, because prefill on a cache hit skips most of the work entirely.

The asymmetry is now inside the models, not just the prices

Until recently you could argue the input and output gap was a pricing convention. DeepSeek's V4.1-Flash makes it structural. It is a 552B parameter mixture-of-experts model that activates 8B parameters for input and 16B for output. The previous generation used 13B for both.

Read that as a design decision rather than a specification. Prefill is cheap per token, so you can afford a leaner pass over the prompt. Decode is expensive per token, so that is where you spend capacity to make each token better. The architecture now encodes the same asymmetry the billing always had.

What this changes about how you write prompts

The instinct most people have is to shorten their prompts. That is usually the wrong lever. A 3,000 token prompt that produces a 200 token answer is already cheap. A 200 token prompt that produces 3,000 tokens of prose is the one running up your bill.

  • Ask for the shape you need. "Return the three IDs" costs a fraction of "explain your reasoning for each candidate and then give the IDs", and for most production paths the IDs are all you consume.

  • Set max_tokens deliberately. Not as a safety net at 4096, but at the length a good answer actually is. It caps the expensive half of the call.

  • Move reasoning out of the output where you can. Some reasoning is worth paying for. Reasoning you immediately discard is pure cost.

  • Put stable context first and reuse it. Prompt caching makes repeated input nearly free, which widens the gap between the two halves even further.

  • Watch for list-and-explain patterns. Asking for 40 items with a sentence each is a long generation wearing a short question's clothes.

None of that means treating input as free. Long context has its own costs in latency and memory, covered in why long context costs more, and KV cache memory is held for the whole generation. But if you are hunting for spend, start at the output end.

FAQ

Is the input and output ratio the same for every model?

No. Three to six times is the common range today, but it moves with architecture and with what the vendor is trying to encourage. Reasoning models, which generate a lot of intermediate tokens you never see, tend to sit at the wider end.

Do reasoning tokens count as output?

Usually yes, and they are billed at the output rate even when they are hidden from you. That is why turning reasoning effort up is expensive in a way that is easy to miss. What is reasoning effort in AI goes into where the setting earns its cost.

Does cached input change the picture?

It widens it. A cache hit removes most of the prefill work, so your input cost drops sharply while output cost stays exactly where it was. Systems with long stable prompts and short answers benefit enormously. See what is prompt caching.

Why does output also feel slower, not just cost more?

Same mechanism. Sequential generation is why you watch answers arrive word by word while the prompt was consumed instantly. Time to first token measures prefill, and everything after it measures decode. What is time to first token splits the two.

For a broader view of what drives model pricing, see why bigger AI models cost more per token. For the layer underneath all of it, how AI models work covers the forward pass this article takes for granted.

How did this land?

About the author

Cecilia Iona
Cecilia Iona

Senior Editor, AI & Product

Cecilia leads the Swarmz editorial desk. She has spent a decade turning complex AI and product topics into writing people actually finish, and she owns the blog's quality bar.

Share

Get the next post in your inbox

One email a month. Product updates, engineering posts, and the best of Built with Swarmz.

I agree to receive emails about AI building tips and Swarmz product news. Unsubscribe any time.