Dashboard

Claude Opus 5.5: Three Price Cuts, Two Catches

Claude Opus 5.5 landed on 22 September with a 20% rate cut, a 60% cache cut and fewer tokens per task. Here is what actually changes on your bill.

Cecilia Iona
Cecilia Iona
Senior Editor, AI & Product
23 September 20261 min read

Anthropic released Claude Opus 5.5 on 22 September 2026 and led with a single number: about 40% cheaper to run than Opus 5 on typical workloads. That number is real, but it is not a price cut. It is three separate changes stacked on top of each other, and only one of them shows up on the pricing page. If you are planning a budget around it, you need to know which is which, because they do not apply evenly to every workload.

What the pricing page says

The published rates for claude opus 5.5 are $4 per million input tokens and $20 per million output tokens, against $5 and $25 for Opus 5. That is a flat 20% cut in both directions.

Rate

Opus 5.5

Opus 5

Change

Input

$4 / M

$5 / M

-20%

Output

$20 / M

$25 / M

-20%

Cache read

$0.20 / M

$0.50 / M

-60%

Cache write

$5 / M

$6.25 / M

-20%

The cache read line is the one worth staring at. A 60% cut on cache reads is three times the headline discount, and it only reaches you if your workload actually reuses a cached prefix. Agent loops that replay a long system prompt and a growing conversation on every turn are mostly cache reads by volume, so they see something close to the full 60%. A one-shot classification call with a short prompt sees none of it. If you have never looked at what fraction of your input tokens are cache hits, this release is the reason to look. Our explainer on how prompt caching works covers where the boundary sits.

The Opus 5.5 discount that is not on the page

The remaining gap between a 20% rate cut and a 40% cost drop comes from token count, not token price. Anthropic says Opus 5.5 uses fewer tokens per task and generates output more than 30% faster than Opus 5. In the company's own testing across GitHub Copilot CLI and VS Code, the model was among the fewest tokens and steps measured, and in VS Code it solved more terminal tasks than Opus 5 in under half the steps.

Half the steps on an agentic task is a much bigger lever than 20% off the rate, and it is also the least portable claim in the release. Steps saved depend on the task. A model that finds the right file on the first read saves you an entire round trip; on a task where it was already finding the right file, it saves you nothing. This is the difference between price per token and cost per task, which we unpack separately in what token efficiency actually means.

The benchmark numbers Anthropic published move in the same direction. Terminal-Bench 4.0 goes from 52.3% to 66.4%, CursorBench 4.0 from 46.6% to 57.8%, OSWorld 2.0 from 74.0% to 81.8%. Those are agentic and computer-use suites, which rewards finishing in fewer attempts rather than answering single questions better.

The two things that got more expensive

First, fast mode. Opus 5.5 offers up to 2.5x the speed in Claude Code and on the Claude Platform at $8 per million input and $40 per million output, which is double the standard rate and 60% above what Opus 5 cost. Speed is now a separate line item rather than a property of the model you picked, so any latency-sensitive path you move onto fast mode reverses the discount and then some.

Second, thinking is no longer optional. Anthropic states that thinking mode is no longer available with thinking switched off. If you were running Opus 5 with thinking disabled to keep short calls cheap and predictable, that configuration does not carry over, and those calls will now produce reasoning tokens you are billed for. This is the failure mode where a cheaper model raises a specific bill: the rate went down 20%, the output token count went up, and the net is worse for that one path. Our piece on reasoning effort covers what the extra tokens buy you and when they are wasted.

What else shipped with it

The release carries safeguards Anthropic describes as similar to those on Claude Fable 5.1 for cybersecurity, biology and distillation, including a preserved thinking anti-distillation measure active for new API accounts, and watermarking for EU AI Act compliance. The model runs on AWS, Google Cloud and Microsoft Azure, and the API string is claude-opus-5-5. Anthropic positions it as performing at the level of Claude Fable 5.1 on most work.

What to do with this

Run your own numbers before you assume the 40%:

  1. Pull one week of usage and split it into input, output, cache read and cache write tokens. The cache split is the part most teams have never measured.

  2. Reprice that week at $4, $20, $0.20 and $5. That is the rate-only saving, which for a cache-heavy workload already exceeds 20%.

  3. Separately, list any call path you ran with thinking off. Those get repriced upward, not downward.

  4. Only then test whether the fewer-steps claim holds on your tasks. It is the largest possible saving and the least likely to transfer unchanged.

Both labs announced cheaper models the same day, which CNBC reported as a response to open-weight competition, and TechCrunch covered alongside the Fable-level performance claim. Two price cuts in one morning is a good moment to re-derive cost per task rather than cost per token, and to revisit our checklist on telling a significant model release from a routine one and our guide to following AI news without drowning in it.

Update, 29 September 2026: Anthropic has since released Claude Sonnet 5.5 at unchanged $2 and $10 list prices with a similar fewer-tokens claim. What Sonnet 5.5 changed covers the numbers.

How did this land?

About the author

Cecilia Iona
Cecilia Iona

Senior Editor, AI & Product

Cecilia leads the Swarmz editorial desk. She has spent a decade turning complex AI and product topics into writing people actually finish, and she owns the blog's quality bar.

Share

Get the next post in your inbox

One email a month. Product updates, engineering posts, and the best of Built with Swarmz.

I agree to receive emails about AI building tips and Swarmz product news. Unsubscribe any time.