Dashboard

What Is Barge-In in Voice AI? Interruptions Explained

Barge-in is the feature that lets someone talk over your voice agent. Getting it wrong is the fastest way to make a voice product feel hostile.

Cecilia Iona
Cecilia Iona
Senior Editor, AI & Product
24 September 20261 min read

Barge-in is the ability to interrupt a voice agent while it is still talking, and have it stop and listen. Say "no, the other one" halfway through a list of options and a system with barge-in cuts itself off mid-word. A system without it keeps reading the list to the end while you repeat yourself, which is the single most reliable way to make a caller hang up.

The name is borrowed from the old phone menu world, where barge-in meant pressing a key before the prompt finished. Voice AI inherited the term and a much harder version of the problem, because now the thing being interrupted is generating speech continuously and the interruption arrives as sound rather than as a keypress.

What has to happen in about 200 milliseconds

Four things, in order, every time someone speaks while the agent is talking:

  1. The system hears incoming audio on the microphone.

  2. It decides whether that audio is a person speaking, rather than a cough, a door, or its own voice arriving back through the speaker.

  3. It decides whether this particular speech should stop the agent.

  4. It halts playback, discards the rest of the generated response, and hands the new audio to the model.

Step four is the one teams forget. Stopping the sound is not enough. If the agent has already generated three sentences and only played one, those other two are sitting in a buffer, and unless they are explicitly discarded the agent will resume reading them after the interruption, answering a question nobody asked any more. That single bug accounts for a large share of voice agents that feel possessed.

The three failure modes

Every barge-in implementation is a compromise between these, and you cannot minimise all three at once.

Failure

What the caller experiences

Usual cause

False trigger

Agent keeps stopping for no reason

Threshold too sensitive, background noise, echo

Missed interruption

Agent talks over the caller

Threshold too strict, or no barge-in at all

Sluggish cutoff

Agent takes a beat too long to stop

Detection window too long, buffered audio not flushed

**False triggers** are the most common and the most damaging, because the agent appears to be listening to a television. The prime suspect is almost never the caller: it is the agent's own output looping back through a speaker into the microphone. Acoustic echo cancellation exists for exactly this, and on a phone line or a laptop without a headset it is not optional. The second suspect is backchannel noise, the "mhm" and "right" that people emit while listening without any intention of taking the floor. A human speaker reads those as encouragement and keeps going. A naive detector reads them as an interruption and stops dead.

**Missed interruptions** usually come from a detector tuned defensively after a false-trigger incident. This is the overcorrection to watch for, because the symptom shows up in call recordings rather than in error logs, and nobody files a bug for it.

**Sluggish cutoff** is a latency problem wearing a different hat. Most detectors require a minimum duration of speech before they commit, typically 100 to 300 milliseconds, precisely so that a cough does not trigger them. Longer windows mean fewer false triggers and a slower stop. That dial is the tradeoff, made explicit.

Semantic barge-in, and why it is harder than it sounds

The sophisticated version does not ask "is someone speaking" but "does what they are saying mean they want the floor". It lets "mhm" and "yeah" through without stopping, and treats "wait" or "no, actually" as a real interruption.

This requires transcribing the interrupting audio far enough to classify it, which costs time you were trying to save, and it fails in the expensive direction: misclassify a genuine interruption as backchannel and you talk over a person who is trying to correct you. Most production systems run the cheap acoustic detector for the stop decision and use the semantic layer only to decide whether to resume where they left off.

Native speech-to-speech models change this picture somewhat, since the model itself hears the interruption rather than a separate detector, but the buffered-audio discard problem is identical and still yours to handle.

Tuning it without a lab

You do not need a research budget, you need recordings of real calls. Pull twenty, and count only two things: how many times the agent stopped when it should not have, and how many times it kept going when it should not have. That ratio tells you which direction to move the threshold, and it is a more useful signal than any synthetic test, because real callers have televisions and children and speakerphones.

Three settings do most of the work. The minimum speech duration before a stop commits. Whether echo cancellation is genuinely on, which is worth verifying rather than assuming. And whether the generated-but-unplayed audio buffer is flushed on interrupt, which is a code path rather than a dial and is the first thing to check when an agent seems to answer questions from the past.

How the agent behaves after an interruption is a prompting decision as much as an engineering one, and prompting a voice agent differs from prompting for text in ways that show up exactly here.

Questions

Is barge-in the same as voice activity detection?

No, though one uses the other. Voice activity detection decides whether the incoming audio contains speech. Barge-in is the whole behaviour: detect, decide, stop playback, flush the buffer, and hand over. Voice activity detection is step two of four.

Should barge-in always be on?

Almost always, with narrow exceptions. Legal disclosures, recording notices and safety warnings are usually configured as non-interruptible, because a caller must be able to say they heard them. Everything else should yield.

Why does the agent stop when nobody is talking?

Echo first, background noise second. Check that acoustic echo cancellation is active on the call path, then listen to a recording of a false trigger. The cause is audible in nine cases out of ten, and it is frequently the agent hearing itself.

Does barge-in add latency?

Slightly, and in the right place. The detection window sits between the caller starting to speak and the agent stopping, so it is latency on interruption rather than on response. Response latency is the separate problem covered by time to first token.

Do I need it for a voice feature inside an app?

If the app speaks in sentences rather than words, yes. Push-to-talk designs can avoid it, at the cost of making the interaction feel like a radio. When adding voice input to an existing app, decide this early, because retrofitting a buffer flush into a playback pipeline is unpleasant.

The underlying question of how any of this gets from sound to meaning is covered in how AI models work.

How did this land?

About the author

Cecilia Iona
Cecilia Iona

Senior Editor, AI & Product

Cecilia leads the Swarmz editorial desk. She has spent a decade turning complex AI and product topics into writing people actually finish, and she owns the blog's quality bar.

Share

Get the next post in your inbox

One email a month. Product updates, engineering posts, and the best of Built with Swarmz.

I agree to receive emails about AI building tips and Swarmz product news. Unsubscribe any time.