Qwen Audio 3.1 Pricing: Voice APIs Cut Up to 95%
Alibaba shipped five speech models and cut audio API prices by up to 95 percent. The interesting part is which tier fell furthest.
Alibaba announced Qwen Audio 3.1 on 23 September 2026 at its Apsara Conference in Hangzhou, shipping five speech models and cutting audio API prices by up to 95 percent. The cuts are not uniform, and the pattern is the most informative thing in the announcement.
What was announced
The release covers speech recognition, text to speech and realtime interaction, adding TTS-Next and ASR-Next alongside the existing ASR, TTS and Realtime endpoints. Reported cuts, per coverage of the announcement:
Capability | Reported cut |
|---|---|
Speech recognition (ASR) | Up to 95 percent |
Text to speech (TTS) | Around 70 percent |
Realtime voice | Around 85 percent |
Separate reporting on the Apsara announcements describes the same lineup alongside latency work on the LiveTranslate product. These are percentage reductions against Alibaba's own previous rates rather than published comparisons with other vendors, which is worth holding onto before treating them as a market ranking.
The pattern matters more than the percentages
Read the three numbers together and they say something specific about where the value has moved.
Speech recognition fell furthest. That is the most commoditised piece of the voice stack: accuracy across major vendors has converged, the task is well defined, and there is little left to differentiate on except price. A 95 percent cut is what a capability looks like when it stops being a product and becomes a utility.
Text to speech fell least. That tracks, because voice quality, cloning and pronunciation control are still genuinely differentiated, and customers pay for a voice they like in a way they do not pay for a marginally better transcript.
Realtime fell hard, at around 85 percent, and that is the number worth watching. Realtime has been the expensive tier that kept native voice agents in pilot rather than production. Cuts of this size move conversational voice from a line item you have to justify into one you can absorb, which is a different kind of change from the other two.
What this changes for a build
Two things genuinely shift.
Workloads that were priced out become sensible. Transcribing every support call rather than a sample, running speech recognition over an entire back catalogue, or leaving a voice agent available on a low-traffic line all change character when the per-minute cost drops by an order of magnitude. The threshold question of what to transcribe stops being an economic one and becomes a privacy and retention one, which is a better problem to have and a harder one to answer.
Pipeline builds get cheaper faster than native ones. If you run separate recognition and synthesis stages, you just got the largest cut on one stage and a modest one on the other. The architectural choice between that and a single native model is worked through in what a speech-to-speech model is, and this pricing shifts the arithmetic toward the pipeline for anything that is not latency-critical.
Four things it does not change
**Latency.** Price and speed are separate levers, and nothing in this announcement claims a faster first response. A voice agent that felt sluggish yesterday feels equally sluggish at a tenth of the cost.
**Interruption handling.** The behaviour that makes a voice agent tolerable is your engineering, not the vendor's rate card. Barge-in is where most voice products actually fail.
**Accuracy on your audio.** Published benchmarks are not your call recordings, with your accents, your product names and your background noise. Measure word error rate yourself on a sample you control before switching.
**Where the audio goes.** A cheaper endpoint is still an endpoint in someone else's data centre, and voice data carries consent and retention obligations that a price cut does not touch.
The wider context
This is the third significant price cut in a week, following GPT-6 Sol and Luna and Claude Opus 5.5. Three vendors cutting within days of each other is a competitive pattern rather than three coincidences, and the structural reasons behind it are covered in why AI API prices keep falling.
For anyone with a voice feature already shipped, the practical move this week is small: re-price your existing usage at the new rates, check whether a workload you shelved on cost grounds is now viable, and resist re-architecting on the strength of a rate card. For anyone about to start, the cost floor for adding voice input to an app is meaningfully lower than it was last week.
How did this land?
About the author

Senior Editor, AI & Product
Cecilia leads the Swarmz editorial desk. She has spent a decade turning complex AI and product topics into writing people actually finish, and she owns the blog's quality bar.


