What Is Token Efficiency in AI Models?
Token efficiency is how many tokens a model burns to finish a task. It explains why the cheaper model per token is often the expensive one per job.
Token efficiency in AI models is how many tokens a model consumes to finish a given task, as opposed to what each of those tokens costs. It is the second half of a cost calculation that most teams only do the first half of, and it is the reason a model advertised at half the price can arrive at the end of the month having charged you more.
The distinction became hard to ignore on 22 September 2026, when Anthropic and OpenAI both cut prices. Anthropic's rate cut on Claude Opus 5.5 was 20%, but the company claimed roughly 40% lower cost on typical workloads. The missing 20 points did not come from the price list. They came from the model using fewer tokens to reach the same answer.
Price per token and cost per task are different numbers
A rate card gives you dollars per million tokens. Your invoice is dollars per million tokens multiplied by the number of tokens you actually used. Only the first factor is published. The second depends on the model, the task and how many attempts it takes.
Here is what that looks like when the two factors point in opposite directions. The rates below are the published September 2026 rates for two real models; the token counts are an illustration of a multi-step coding task, not a measured benchmark.
Model A ($4 / $20 per M) | Model B ($2 / $10 per M) | |
|---|---|---|
Input tokens used | 40,000 | 110,000 |
Output tokens used | 6,000 | 20,000 |
Input cost | $0.16 | $0.22 |
Output cost | $0.12 | $0.20 |
Cost per task | $0.28 | $0.42 |
Model B is half the price per token and 50% more expensive per task. Nothing unusual has happened. Model B read more files before it found the right one, made two attempts where Model A made one, and wrote more words in each reply. Those are ordinary differences between models, and on an agentic workload they dominate the rate.
Where the tokens go
Four things drive token count on the same nominal task.
Steps. An agent that locates the relevant function on its first read does one round trip. One that reads six files first does six, and every one of those reads is billed as input. Anthropic reported that in VS Code testing, Opus 5.5 solved more terminal tasks than Opus 5 in less than half the steps. Halving steps on a loop that replays context each turn is a larger saving than any realistic rate cut.
Context replay. In a multi-turn agent, the conversation so far is resent on every turn. Ten turns of a growing transcript means the early turns are paid for ten times. This is why the same task costs wildly different amounts depending on how the loop is written rather than which model runs it.
Reasoning tokens. Models that think before answering bill you for the thinking. More reasoning usually means fewer retries, so it is not waste by default, but it is a real cost that does not appear in the answer. We cover the trade in reasoning effort.
Verbosity. Output tokens cost several times what input tokens cost on every major model. A model that returns the patch costs less than one that returns the patch with three paragraphs of explanation you did not ask for.
How to measure it on your own workload
Vendor efficiency claims are averages over their task mix, not yours. The measurement is not difficult.
Pick ten real tasks from your logs that you can grade pass or fail without arguing about it. Ten is enough to see a factor-of-two difference, which is the size of difference that matters.
Run each task on each candidate model with the same loop, the same tools and the same prompt. Change one variable at a time or you will learn nothing.
Record four numbers per run: input tokens, output tokens, number of model calls, and pass or fail. The call count is the one people forget and the one that explains the rest.
Compute total cost divided by tasks that passed. Not cost per run. A model that fails a third of the time has its cost per success inflated by half, and that is the number you are actually buying.
The output is a dollar figure per completed task, which is the only figure you can put in a budget. Our guide to estimating tokens before you run anything covers the sizing arithmetic if you need a number before you have logs.
Why token efficiency and capability are converging
Token efficiency used to be a property of small models. It is increasingly a property of good ones. A model that understands a codebase well enough to go straight to the right file uses fewer tokens because it is more capable, not less. That is why the September benchmarks that moved most on Opus 5.5 were agentic suites like Terminal-Bench and OSWorld rather than single-answer tests: they reward finishing in fewer moves.
The older intuition, that spending more per token buys you more, still holds for the reasons set out in why bigger models cost more per token. What has changed is that the cheaper model is no longer automatically the cheaper choice, because capability now shows up as a discount on the other factor.
A recent worked example: Sonnet 5.5 vs GPT-6.1 Sol shows two models with identical list prices ending up with different bills.
FAQ
Is token efficiency the same as speed?
No, though they are related. Speed is tokens per second, which affects how long a user waits. Efficiency is tokens per task, which affects what you pay. A model can be fast and wasteful, producing a great deal of unnecessary output very quickly.
Does a bigger context window improve token efficiency?
Not on its own. A larger window lets you put more into a single call, which can remove round trips, but it also makes it easy to send far more context than the task needs and pay for all of it. The window is capacity, not a discount.
How much can token efficiency vary between models on the same task?
Enough to reverse a price comparison. Differences of two to three times in tokens per completed task are common on multi-step agentic work, against rate differences that are usually well under two times. On single-turn tasks with short prompts the variation is much smaller and the rate card is a reasonable guide.
Should I optimise for efficiency or just pick the cheapest model?
Measure before you pick. On short, well-defined, high-volume calls the cheapest model is usually genuinely cheapest. On long agentic tasks the ranking frequently inverts, and the only way to know which regime you are in is to run the ten-task comparison above. For the underlying mechanics of what a token is and how models consume them, start with what is a token in AI and how AI models work.
How did this land?
About the author

Senior Editor, AI & Product
Cecilia leads the Swarmz editorial desk. She has spent a decade turning complex AI and product topics into writing people actually finish, and she owns the blog's quality bar.


