Why Do AI API Prices Keep Falling? Four Reasons
Three major price cuts landed in a single week. The reasons are structural, and they tell you something useful about how to architect around them.
In the week to 24 September 2026, three separate vendors cut prices. OpenAI's GPT-6 tier landed at roughly half the previous generation's rate, Anthropic's Opus 5.5 came in around 40 percent below Opus 5, and Alibaba cut its audio APIs by up to 95 percent. That is not a promotion. It is the normal operating rhythm of this market, and it has been for three years.
The short answer to why: the cost of serving a given level of capability falls faster than the capability improves, and competition forces vendors to pass it on. The longer answer is four distinct mechanisms, which matter because they fall at different rates and in different places.
1. The same answer gets cheaper to compute
Most of the drop has nothing to do with new hardware. It comes from serving improvements on models that already exist: better batching, so one pass over the weights serves many requests at once; caching of repeated prompt prefixes; speculative decoding, where a small model drafts and a large one verifies. None of these change what the model says. They change how many requests one machine can serve per second, and that number goes almost directly into the price.
2. Smaller models catch up to last year's large ones
Distillation and better training data mean the model you need for a given job keeps shrinking. A task that required a frontier model in 2024 is routinely handled by something a fraction of the size now, and a smaller model is cheaper to run at every step. Vendors are not only cutting the price of a fixed thing, they are moving the floor of what counts as good enough.
3. Competition with a very short memory
Price is one of the few axes where a switch is genuinely easy. Swapping vendors is a config change and an evaluation run, not a migration. That makes published rates strategically load-bearing in a way they are not in most infrastructure markets, and it is why cuts arrive in clusters within days of each other, as they did this week.
4. Capacity gets built ahead of demand
Serving capacity is bought in large, lumpy increments. When a vendor brings a new tranche online it is cheaper per unit than the last one and it is idle until filled. Cutting price is how you fill it. This is the mechanism that occasionally reverses, which is worth remembering before you build a business on the trend continuing in a straight line.
The asymmetry that should change your architecture
Here is the part that rarely makes the announcement summaries: input and output tokens do not fall at the same rate. Input is consistently cheaper than output across every major vendor, often by a factor of three to five, and the gap has been widening rather than closing.
The reason is mechanical. Reading a prompt is one parallel pass over all of it. Writing a response is one full pass over the model weights per token produced, strictly in sequence. The first parallelises beautifully and gets cheaper with every batching improvement. The second is bounded by how fast memory can be read, and that has improved far more slowly.
The design consequence is concrete:
Pattern | Cost behaviour | Verdict |
|---|---|---|
Long prompt, short structured answer | Cheap and getting cheaper | Lean into it |
Short prompt, long generated prose | Expensive, falling slowly | Question it |
Long prompt repeated across requests | Nearly free with caching | Structure for it |
Chain of agent steps, each generating | Output cost multiplies per hop | Audit it |
If you have been rationing context to save money, you are optimising the wrong side of the bill. Give the model more to read and ask it to write less. A prompt that includes the whole document and requests a twelve-field JSON object is usually cheaper than a terse prompt that asks for three paragraphs of explanation, and it is generally more useful too. The arithmetic for your own case is in how much an AI feature costs per user.
What to do about a cut when it lands
Falling prices are only good news if you notice them. Three habits cover it.
**Re-run your model choice quarterly, not once.** The cheapest model that passes your evaluation set changes every few months. If you picked a model in March and have not revisited it, you are probably overpaying for capability you tested once. Pinning a model version protects you from surprise behaviour changes, and it also freezes you out of price improvements unless somebody reviews the pin.
**Decide in advance what a cut does to your pricing.** Whether the saving reaches your customer is a positioning question you want to have answered before a client reads the news and asks. Whether to pass an AI price cut to your client works through both positions.
**Recheck the build versus buy line.** Every API price cut moves the threshold at which running your own model makes sense, and it moves it against self-hosting. The comparison is worked through in whether self-hosting an open weight model is cheaper.
What falling prices do not fix
Three things stay stubbornly where they are. Latency has improved far less than cost, so a cheaper model is not a faster one. Accuracy on your specific task is unrelated to price, and a cheaper model that is wrong more often is not a saving. And per-token price is not per-task cost: an agent that retries three times at half price costs more than one that succeeds at full price.
The one genuine risk in this trend is planning around it. Rates have fallen consistently, but they are a vendor decision rather than a law of physics, and capacity economics can reverse. Budget on today's published rate. Treat the next cut as upside rather than as a line in the forecast.
Questions
Will AI API prices keep falling?
The mechanisms driving the drop are still active, so the direction is likely. The rate is not predictable, and a vendor under margin pressure can raise prices or quietly retire a cheap tier. Plan on current rates.
Why is output so much more expensive than input?
Reading your prompt is a single parallel pass over the whole thing. Generating output is one sequential pass over the model weights per token. The first benefits from every batching improvement, the second is limited by memory bandwidth and improves slowly.
Should I switch vendors every time one cuts prices?
No. Switching costs an evaluation cycle and some risk, and the saving on a small workload rarely covers it. Review on a schedule instead, and switch when the gap is large enough to matter at your actual volume.
Do cheaper models mean worse results?
Not automatically. Price tracks the cost of serving, which tracks model size, which correlates with capability but does not determine it on your specific task. Measure on your own evaluation set, since that is the only comparison that applies to you.
How should I bill clients when my costs keep dropping?
Avoid tying a client price directly to a vendor rate you do not control in either direction. The structures that survive vendor volatility are covered across the AI monetization strategies guide.
How did this land?
About the author

Growth & SEO Lead
Manuele covers distribution: SEO, content strategy, and how AI-built products find their first thousand users. He tests everything he recommends.


