Dashboard

Why Do AI API Prices Keep Falling? Four Reasons

Three major price cuts landed in a single week. The reasons are structural, and they tell you something useful about how to architect around them.

Manuele Estivo
Manuele Estivo
Growth & SEO Lead
24 September 20261 min read

In the week to 24 September 2026, three separate vendors cut prices. OpenAI's GPT-6 tier landed at roughly half the previous generation's rate, Anthropic's Opus 5.5 came in around 40 percent below Opus 5, and Alibaba cut its audio APIs by up to 95 percent. That is not a promotion. It is the normal operating rhythm of this market, and it has been for three years.

The short answer to why: the cost of serving a given level of capability falls faster than the capability improves, and competition forces vendors to pass it on. The longer answer is four distinct mechanisms, which matter because they fall at different rates and in different places.

1. The same answer gets cheaper to compute

Most of the drop has nothing to do with new hardware. It comes from serving improvements on models that already exist: better batching, so one pass over the weights serves many requests at once; caching of repeated prompt prefixes; speculative decoding, where a small model drafts and a large one verifies. None of these change what the model says. They change how many requests one machine can serve per second, and that number goes almost directly into the price.

2. Smaller models catch up to last year's large ones

Distillation and better training data mean the model you need for a given job keeps shrinking. A task that required a frontier model in 2024 is routinely handled by something a fraction of the size now, and a smaller model is cheaper to run at every step. Vendors are not only cutting the price of a fixed thing, they are moving the floor of what counts as good enough.

3. Competition with a very short memory

Price is one of the few axes where a switch is genuinely easy. Swapping vendors is a config change and an evaluation run, not a migration. That makes published rates strategically load-bearing in a way they are not in most infrastructure markets, and it is why cuts arrive in clusters within days of each other, as they did this week.

4. Capacity gets built ahead of demand

Serving capacity is bought in large, lumpy increments. When a vendor brings a new tranche online it is cheaper per unit than the last one and it is idle until filled. Cutting price is how you fill it. This is the mechanism that occasionally reverses, which is worth remembering before you build a business on the trend continuing in a straight line.

The asymmetry that should change your architecture

Here is the part that rarely makes the announcement summaries: input and output tokens do not fall at the same rate. Input is consistently cheaper than output across every major vendor, often by a factor of three to five, and the gap has been widening rather than closing.

The reason is mechanical. Reading a prompt is one parallel pass over all of it. Writing a response is one full pass over the model weights per token produced, strictly in sequence. The first parallelises beautifully and gets cheaper with every batching improvement. The second is bounded by how fast memory can be read, and that has improved far more slowly.

The design consequence is concrete:

Pattern

Cost behaviour

Verdict

Long prompt, short structured answer

Cheap and getting cheaper

Lean into it

Short prompt, long generated prose

Expensive, falling slowly

Question it

Long prompt repeated across requests

Nearly free with caching

Structure for it

Chain of agent steps, each generating

Output cost multiplies per hop

Audit it

If you have been rationing context to save money, you are optimising the wrong side of the bill. Give the model more to read and ask it to write less. A prompt that includes the whole document and requests a twelve-field JSON object is usually cheaper than a terse prompt that asks for three paragraphs of explanation, and it is generally more useful too. The arithmetic for your own case is in how much an AI feature costs per user.

What to do about a cut when it lands

Falling prices are only good news if you notice them. Three habits cover it.

  • **Re-run your model choice quarterly, not once.** The cheapest model that passes your evaluation set changes every few months. If you picked a model in March and have not revisited it, you are probably overpaying for capability you tested once. Pinning a model version protects you from surprise behaviour changes, and it also freezes you out of price improvements unless somebody reviews the pin.

  • **Decide in advance what a cut does to your pricing.** Whether the saving reaches your customer is a positioning question you want to have answered before a client reads the news and asks. Whether to pass an AI price cut to your client works through both positions.

  • **Recheck the build versus buy line.** Every API price cut moves the threshold at which running your own model makes sense, and it moves it against self-hosting. The comparison is worked through in whether self-hosting an open weight model is cheaper.

What falling prices do not fix

Three things stay stubbornly where they are. Latency has improved far less than cost, so a cheaper model is not a faster one. Accuracy on your specific task is unrelated to price, and a cheaper model that is wrong more often is not a saving. And per-token price is not per-task cost: an agent that retries three times at half price costs more than one that succeeds at full price.

The one genuine risk in this trend is planning around it. Rates have fallen consistently, but they are a vendor decision rather than a law of physics, and capacity economics can reverse. Budget on today's published rate. Treat the next cut as upside rather than as a line in the forecast.

Questions

Will AI API prices keep falling?

The mechanisms driving the drop are still active, so the direction is likely. The rate is not predictable, and a vendor under margin pressure can raise prices or quietly retire a cheap tier. Plan on current rates.

Why is output so much more expensive than input?

Reading your prompt is a single parallel pass over the whole thing. Generating output is one sequential pass over the model weights per token. The first benefits from every batching improvement, the second is limited by memory bandwidth and improves slowly.

Should I switch vendors every time one cuts prices?

No. Switching costs an evaluation cycle and some risk, and the saving on a small workload rarely covers it. Review on a schedule instead, and switch when the gap is large enough to matter at your actual volume.

Do cheaper models mean worse results?

Not automatically. Price tracks the cost of serving, which tracks model size, which correlates with capability but does not determine it on your specific task. Measure on your own evaluation set, since that is the only comparison that applies to you.

How should I bill clients when my costs keep dropping?

Avoid tying a client price directly to a vendor rate you do not control in either direction. The structures that survive vendor volatility are covered across the AI monetization strategies guide.

How did this land?

About the author

Manuele Estivo
Manuele Estivo

Growth & SEO Lead

Manuele covers distribution: SEO, content strategy, and how AI-built products find their first thousand users. He tests everything he recommends.

Share

Get the next post in your inbox

One email a month. Product updates, engineering posts, and the best of Built with Swarmz.

I agree to receive emails about AI building tips and Swarmz product news. Unsubscribe any time.