What Is Tokens Per Second in AI?
Three different measurements all get called speed, they move independently, and the one on the spec sheet is usually the one that flatters the model.
What Is Tokens Per Second in AI?
What is tokens per second in AI? It is the rate at which a model produces output, measured in the chunks of text it actually generates rather than in words. A model quoted at 80 tokens per second is emitting roughly 60 words per second, which is comfortably faster than anyone reads. It is the number vendors lead with when they talk about speed, and on its own it tells you remarkably little about how fast your application will feel.
The reason is that three different measurements all get called speed, they move independently, and the one you are shown is usually the one that flatters the model.
What is tokens per second in AI actually measuring?
Metric | What it measures | What the user experiences |
|---|---|---|
Time to first token | The wait between sending a request and the first output appearing | How long the screen sits empty |
Output tokens per second | The generation rate once output has started, for one request | How fast text scrolls after it starts |
Total throughput | Tokens per second across every concurrent request on the server | Whether the app slows down when it gets busy |
These come apart in ways that matter. A model can be quick to start and slow to generate, or the reverse. The industry terminology is reasonably settled now, and NVIDIA's inference benchmarking documentation is a clear reference for how the metrics are defined and measured.
Why the number in the benchmark is not the number you get
Published tokens per second figures are almost always single-stream: one request, nothing else running, on hardware configured for the test. Production has concurrency, and concurrency changes the arithmetic completely.
Serving systems batch requests together, because a GPU generating for one user is mostly idle. Batching raises total throughput substantially while lowering the per-request rate. A server doing 80 tokens per second for a single user might do 25 tokens per second each for forty concurrent users, which is 1,000 tokens per second in total and a noticeably slower experience for every individual one of them.
Single-stream speed is a property of the model. The speed your users see is a property of how busy your server is.
This is why the same model can feel fast in a demo and sluggish at four in the afternoon, with no configuration having changed. It is also why comparing a provider's published figure against your own observed latency usually produces a disappointing result that nobody has misreported. The underlying mechanics are covered in how batch inference works.
Why generation is slower than reading the prompt
A long prompt is processed far faster than a short response is generated, which strikes most people as backwards. The reason is that reading the prompt happens in parallel, all at once, while generating output happens one token at a time, each one depending on the one before it. The prefill and decode split is the standard name for this, and it explains a lot of otherwise confusing behaviour, including why output tokens cost more than input tokens on essentially every provider's price list.
What to measure instead
For most applications, tokens per second is the wrong headline number. Measure these three things on your own traffic instead.
Time to first token at the 95th percentile, not the average. The average hides the slow requests, and the slow requests are the ones people complain about.
End-to-end time for a complete typical response. Users care when the answer is finished, not about the rate it arrived at.
How both of those degrade at your busiest hour compared with your quietest. If the gap is large, your problem is capacity, not model choice.
If the screen sitting empty is the complaint, generation rate is not your bottleneck and a faster model will not fix it. That is a time to first token problem, and it usually comes from prompt size, cold starts or queueing rather than from the model's raw speed.
Rough figures for calibration
These are orders of magnitude rather than promises, and they move with every release, but they help when reading a spec sheet.
Comfortable reading speed is roughly 5 to 8 tokens per second. Anything above that outpaces the reader.
Streaming output above about 20 tokens per second reads as instant to most people.
Conversational use tolerates a second or so before the first token. Much beyond that and the interface feels broken rather than slow.
Agent work that runs unattended does not care about generation rate at all. Total throughput and cost are the only figures that matter there.
That last point is worth sitting with, because it inverts the usual advice. If a model is producing output nobody is watching, optimising for perceived speed is wasted effort. Optimise for throughput per unit of cost and let individual requests take as long as they take.
Frequently asked questions
How many words is a token?
Around 0.75 words for ordinary English, so 100 tokens is roughly 75 words. Code, other languages and unusual names all shift the ratio, sometimes considerably. What a token is covers why the unit is not a word in the first place.
Do faster models produce worse answers?
Not inherently, but the fast variants in a family are usually smaller or less computation-heavy, and that trade-off is real. The right comparison is quality per unit of time on your own tasks, not speed alone.
Why does the same model report different speeds on different providers?
Hardware, batching configuration, quantisation, and how loaded the service is. Identical open weights served by two providers can differ by a factor of several, which is why benchmarks pin the serving setup as well as the model.
Can I do anything about it from my side?
Some. Shorter prompts reduce time to first token, streaming makes the wait feel shorter without changing it, and caching a repeated prefix helps considerably. None of that changes the model's generation rate, but all of it changes what the user experiences, which is the thing you were actually trying to improve.
How did this land?
About the author

Senior Editor, AI & Product
Cecilia leads the Swarmz editorial desk. She has spent a decade turning complex AI and product topics into writing people actually finish, and she owns the blog's quality bar.


