What Is a Speech-to-Speech Model? Voice AI Explained
Speech-to-speech models take audio in and give audio back, with no text stage in the middle. Here is what that changes, and what it quietly takes away.
A speech-to-speech model takes audio in and produces audio out, with no text stage in between. Ask it a question out loud and it answers out loud, having never written your words down. That is the whole distinction, and it matters for one reason above all others: every stage you remove is latency you stop paying, and voice is the one interface where latency is the product.
The alternative, still the more common build, is a pipeline: a speech recognition model turns your audio into text, a language model reads the text and writes a reply, and a text-to-speech model reads that reply aloud. Three models, three handoffs, three chances to wait.
The latency arithmetic, which is the entire argument
Add up a pipeline turn honestly and the numbers explain the industry's enthusiasm:
Stage | What happens | Rough cost |
|---|---|---|
Endpointing | Deciding the human stopped talking | 200 to 700 ms |
Speech recognition | Audio to text | 100 to 400 ms |
Language model | Reading text, producing the first token | 300 ms to several seconds |
Text to speech | Text to the first audio sample | 100 to 300 ms |
Those stages are mostly sequential, so they add. A pipeline that looks fast in every individual benchmark can still land a full second of silence between a person finishing a sentence and hearing anything back. Human conversation runs on gaps of roughly 200 milliseconds, which is why a one second pause does not read as "thinking", it reads as "broken".
A native speech-to-speech model collapses the middle. It processes audio tokens directly and emits audio tokens directly, so there is one model to wait on rather than three, and it can begin responding before the sentence is fully resolved. The first stage, deciding the person has stopped, does not disappear. Nothing removes that. But the other three fold into one.
The cost side is now moving fast. Alibaba's Qwen Audio 3.1 release on 23 September 2026 cut its realtime voice model price by roughly 85 percent, alongside cuts of up to 95 percent on speech recognition and about 70 percent on text to speech, according to coverage of the Apsara Conference announcement. Realtime audio has been the expensive tier for two years. It is getting less expensive quickly.
What "no text" actually means
The model is not secretly transcribing and hiding it. Audio is turned into tokens the same way text is, just from a different vocabulary: short slices of sound rather than fragments of words. If you have read what a token in AI actually is, the mental model transfers directly. Same machinery, different alphabet.
This is why a native model can hear things a transcript cannot carry. Sarcasm. Hesitation. Someone trailing off because they are unsure rather than because they finished. An accent it can mirror. A transcript flattens all of that into the same flat string of words, and the language model downstream never knows it happened.
The three things you give up
This is the part the launch posts skip.
**The transcript you were logging.** In a pipeline, the text in the middle is free, and you were probably using it for more than you realised: support ticket search, quality review, analytics, dispute evidence. Go native and you either lose it or pay a separate transcription pass to get it back, which reintroduces cost but not latency, since it can run after the fact.
**The moderation hook.** Filtering, redaction, and policy checks are easy to apply to a string and awkward to apply to a waveform. Any guardrail you built on the text step needs rebuilding, and the audio-native equivalents are less mature.
**Voice selection and consistency.** Dedicated text-to-speech vendors compete on voice libraries, cloning, and fine pronunciation control. Native models typically ship a handful of voices and less control over how a specific name or product term is pronounced.
None of these are disqualifying. All of them are cheaper to solve before you build than after.
Which one to pick
Native speech-to-speech earns its place when the conversation is the product and the turns are fast: a booking line, a drive-through, an interruption-heavy support call, anything where a person will talk over the machine. That interruption case has its own mechanics, covered in what barge-in means in a voice agent.
A pipeline stays the right answer when the text in the middle is doing real work. If the model has to call your database, follow a scripted compliance flow, or hand off to a tool, you want a text stage you can inspect, test and version. It is also the easier build, and the one where adding voice input to an app you already have is a bolt-on rather than a rewrite.
A useful middle position exists: run the pipeline, but stream every stage so nothing waits for a complete unit before starting the next. Most of the perceived gap in a naive pipeline comes from waiting for whole sentences rather than from model speed. Streaming recovers a surprising amount of it for none of the architectural cost.
Questions
Is a speech-to-speech model the same as a realtime API?
Mostly yes, in practice. Vendors market native audio models under "realtime" branding because low latency is the selling point. Check whether audio genuinely goes in and out of one model, or whether the vendor is wrapping a fast pipeline behind a streaming interface. Both can be good. Only one removes the handoffs.
Does it still hallucinate?
Yes. It is the same kind of model with a different input and output format. Nothing about processing audio makes a model more truthful, and mishearing a name or a number is now a failure with no transcript to check afterwards.
How do I measure accuracy without a transcript?
You largely cannot use the standard metric directly, since word error rate needs text on both sides. Teams usually run a separate transcription pass over recorded calls purely for evaluation, which is acceptable offline even when it would be too slow in the live path.
Is it cheaper than a pipeline?
Not reliably, though the gap is narrowing fast. Audio tokens are billed at a higher rate than text tokens at most vendors because there are far more of them per second of speech. Price the specific turn length you expect rather than trusting a per-token comparison, and account for whatever you spend re-transcribing for logs.
Do I still need to worry about the first response being slow?
Yes, and more than in a chat interface. The same time to first token pressure applies, except silence in a voice call reads as a dropped connection rather than a slow page. Under roughly 500 milliseconds is the bar, and it is a hard one.
For the wider picture of how any of these models turn input into output in the first place, how AI models work covers the mechanics that all of them share.
How did this land?
About the author

Senior Editor, AI & Product
Cecilia leads the Swarmz editorial desk. She has spent a decade turning complex AI and product topics into writing people actually finish, and she owns the blog's quality bar.


