Dashboard

Sonnet 5.5 vs GPT-6.1 Sol: Same Price, Different Bill

Sonnet 5.5 vs GPT-6.1 Sol: both list at $2 and $10 per million tokens. A worked cost example shows why bills differ and what to test on your own tasks.

Carlo Zuercher
Carlo Zuercher
Staff Engineer, Platform
29 September 20261 min read

Sonnet 5.5 vs GPT-6.1 Sol is not a price fight, because the list prices are identical: $2 per million input tokens and $10 per million output tokens for both. The difference shows up in three smaller places: cache pricing, how many tokens each model spends per task, and what each vendor has actually measured. Below is a worked example, the claims that cannot be compared, and a test that settles it on your own workload.

Sonnet 5.5 vs GPT-6.1 Sol: the list prices

Per the launch coverage of Claude Sonnet 5.5 and GPT-6.1 Sol:

  • Standard input: $2 per million tokens on both

  • Output: $10 per million tokens on both

  • Cached input: $0.20 per million on Sonnet 5.5, $0.10 per million on GPT-6.1 Sol

  • Cache writes: $2.50 per million on Sonnet 5.5. The GPT-6.1 Sol coverage we read gave no separate cache write price.

So on paper, the only price gap is cached input, where Sol is half the price.

A worked example (hypothetical numbers)

The token counts below are invented to show the arithmetic, not measured from either model. Picture an agent task that sends 300,000 input tokens across its turns, of which 240,000 are cache reads on a reused prefix, and produces 15,000 output tokens.

  • Sonnet 5.5: 60,000 fresh input tokens at $2 per million is $0.12. 240,000 cached at $0.20 per million is $0.048. 15,000 output at $10 per million is $0.15. Total: $0.318.

  • GPT-6.1 Sol: the same $0.12 fresh input and $0.15 output, with the cached part at $0.10 per million, which is $0.024. Total: $0.294.

Sol is about 7.5% cheaper on that task from cache pricing alone. Now change one assumption: suppose Sol takes 20,000 output tokens to finish the same job. That adds $0.05 and Sol lands at $0.344, above Sonnet's $0.318. A 5,000 token difference in how much a model writes outweighed the whole cache discount.

The lesson is the one in our explainer on token efficiency: the price sheet sets the rate, and the model decides the quantity. With identical rates, quantity decides the winner.

The claims you cannot line up

Both launches came with numbers, and they do not match up:

  • Anthropic reports Terminal-Bench 4.0 at 70.6% for Sonnet 5.5. OpenAI's reported cost figure for Sol is per task on Terminal-Bench Science, a different benchmark, where it cites $5.47 per task at maximum effort.

  • Each vendor picks the competitor and settings that suit its comparison. OpenAI compares Sol with Opus 5.5 and Astra; Anthropic compares Sonnet 5.5 with Sonnet 5.

  • Anthropic's "up to 30% less per task" is against its own previous model. OpenAI's "a fifth of Astra's price" is against a more expensive tier.

None of these is a Sonnet 5.5 against Sol measurement. If you want that, you have to run it yourself, and how to spot an inflated AI benchmark claim explains what to check in the meantime.

Practical differences that are not about price

  • Availability: Sonnet 5.5 is offered on Amazon Web Services, Google Cloud and Microsoft Azure, with zero data retention. GPT-6.1 Sol is available in the OpenAI API, ChatGPT Work and Codex, and not yet in Chat.

  • Model line: Sonnet sits under Opus 5.5 in Anthropic's lineup, with Haiku 5.5 due in the coming weeks. Sol is the cheaper line that shipped after OpenAI cancelled GPT-6.1 Astra, covered in our report on the cancellation.

A test that settles it in an afternoon

  1. Collect 25 real tasks from your logs, covering your common cases and your worst.

  2. Run each on both models at the effort setting you would ship.

  3. Log fresh input, cached input, output tokens and tool calls for every run.

  4. Price each run with the rates above and compare cost per passing task, not cost per call.

  5. Read the failures by hand.

Whichever wins, keep the harness. Model releases now arrive weekly, and how to know when to upgrade to a newer AI model is easier when the test already exists. For a broader routing view, see which AI model for which task.

Which one to try first

Start with the one that fits your existing setup, since switching costs are real. If your stack already runs through Amazon Web Services, Google Cloud or Azure, Sonnet 5.5 is offered on all three. If you already use Codex or ChatGPT Work, Sol is in both from launch. Then run the test above on the second one. If both pass, take the cheaper cost per passing task and keep the other configured as a fallback, because an outage or a pricing change at one vendor is a matter of when, not if. Route the hardest tasks to a larger tier such as Opus 5.5 only where your own results show the smaller models failing; what model routing is explains the mechanics.

Falling list prices are part of the same picture, and why AI API prices keep falling covers the forces behind them.

FAQ

Is Sonnet 5.5 cheaper than GPT-6.1 Sol?

Their list prices for input and output are the same. Sol has a lower cached input price. Your real cost depends on how many tokens each uses on your tasks.

Which is better for coding, Sonnet 5.5 or GPT-6.1 Sol?

The vendors report different benchmarks, so no like-for-like result exists in the launch coverage. Run both on your own repository and compare cost per passing task.

Does Sonnet 5.5 have zero data retention?

Anthropic offers Sonnet 5.5 with zero data retention, per the launch coverage. Check the current terms for your account.

How did this land?

About the author

Carlo Zuercher
Carlo Zuercher

Staff Engineer, Platform

Carlo works on the platform that turns prompts into running apps. He writes the engineering deep dives and the changelog notes worth reading.

Share

Get the next post in your inbox

One email a month. Product updates, engineering posts, and the best of Built with Swarmz.

I agree to receive emails about AI building tips and Swarmz product news. Unsubscribe any time.