What Is a Stop Sequence in AI?
A string that ends generation the moment the model produces it. What it does, what it silently removes, and when to reach for something else.
What Is a Stop Sequence in AI?
A stop sequence is a string you give a model in advance that tells it to stop generating the moment it produces that string. It is an API parameter, not something the model decides. You pass one or more strings, and generation halts at the first match. Crucially, the matched string is normally not included in what comes back, which is the source of most of the confusion people have with the feature.
What a stop sequence actually does
During generation the model emits one token at a time. After each token the serving layer checks whether the text produced so far ends with any of your stop strings. If it does, generation ends immediately and the response is returned with a finish reason indicating a stop sequence was hit, rather than a natural completion.
Two things follow from that mechanism. First, it is a string match on the output text, not a semantic instruction: the model is not persuaded to stop, it is cut off. Second, it operates after the fact, so the model has no idea it is about to be stopped and does not plan its output around the boundary.
This is a different thing from the end-of-sequence token the model produces on its own when it considers an answer finished, and different again from hitting a token limit. All three end a response, which is why they get conflated. If your responses are ending in the wrong place and you are not sure which of the three is responsible, why AI responses get cut off mid-sentence walks through distinguishing them from the finish reason.
Where stop sequences earn their keep
Use case | Stop sequence | Why |
|---|---|---|
Multi-turn transcript format | "\nUser:" | Stops the model writing the human's next line for them |
Single field extraction | "\n" | One line out, nothing after it |
Structured block | "```" | Ends at the close of a code fence |
Custom delimiter | "<<END>>" | An explicit marker you asked the model to emit |
The transcript case is the original reason the parameter exists. Ask a base model to continue a dialogue and it will cheerfully write both sides of it, including the user's next question. Setting the next speaker's prefix as a stop sequence ends the turn where it should end.
The custom delimiter case is the one worth knowing for modern work. If you ask a model to emit a specific marker when it is done, and set that marker as a stop sequence, you get a reliable boundary you control rather than guessing from the text. That pairs well with the techniques in getting consistent output every time.
The three ways stop sequences go wrong
The delimiter disappears
You set the stop sequence to your closing marker, then write a parser that looks for that marker to know where the payload ends. It is never there, because the API excluded it. Your parser finds nothing and reports an empty result on responses that were actually fine. Either parse without expecting the delimiter, or append it yourself after the call.
It fires inside legitimate content
A stop sequence of a single newline works perfectly until the model produces a legitimately multi-line answer, at which point you silently receive the first line only. The same happens with "```" when the answer contains a code block in the middle rather than at the end. Choose delimiters that cannot plausibly occur in the content, which is why arbitrary markers beat punctuation.
It is treated as an instruction
People sometimes set a stop sequence and assume the model now knows to wrap up neatly before reaching it. It does not. The model was never told. If you want a tidy ending you have to ask for it in the prompt as well, and then use the stop sequence as the enforcement layer. The general point about instructions not constraining generation the way people expect is covered in why telling AI not to do something does not work.
Stop sequence versus max tokens
These solve different problems and you generally want both. A stop sequence gives you a semantic boundary: stop when the content reaches this point. A token limit gives you a cost and latency ceiling: stop after this much output regardless. A response that ends on your stop sequence is complete. A response that ends on the token limit is truncated, and probably unusable.
Check the finish reason on every call if you are doing anything automated with the output. Treating a length-truncated response as a complete one is how half a JSON object gets written to your database. If you need structured output reliably, the newer approach is to constrain the format directly rather than to police the end of it, which is what structured output does, and it makes stop sequences largely unnecessary for that particular job.
Practical defaults
Use an unlikely explicit marker rather than punctuation or whitespace.
Ask for the marker in the prompt and set it as the stop sequence. Belt and braces.
Always read the finish reason. Never assume a response ended where you intended.
Keep the list short. Most APIs cap it at around four, and each additional string is another chance to fire early.
For JSON, prefer a structured output mode if the provider has one.
Stop sequences are a small parameter with an outsized effect on whether automated pipelines behave, and they cost nothing to get right. For the surrounding concepts, how AI models work covers the generation loop this parameter interrupts.
Frequently asked questions
What is a stop sequence in AI?
A string passed to the model's API that ends generation as soon as the output contains it. It is enforced by the serving layer rather than decided by the model, and the matching string is normally excluded from the text you receive.
Is the stop sequence included in the model's response?
Usually not. Most providers cut the output at the point the match begins and return everything before it, which surprises people who wrote a parser expecting to find the delimiter. Check your provider's documentation, and append the marker yourself if downstream code needs it.
What is the difference between a stop sequence and max tokens?
A stop sequence ends generation at a place you defined in the content, so the response is complete. Max tokens ends it after a fixed amount of output regardless of where that falls, so the response may be truncated mid-sentence. The finish reason tells you which happened.
Can I use more than one stop sequence?
Yes, most APIs accept several, commonly up to four, and generation ends at whichever matches first. Keep the list short, because every extra string increases the chance one of them appears legitimately inside the content and ends the response early.
Do I still need stop sequences if I use structured output?
Rarely. A structured output or JSON mode constrains the shape of the whole response, which handles the boundary problem more reliably than matching on a string. Stop sequences remain useful for free-text formats such as transcripts and for custom markers in plain prose.
How did this land?
About the author

Senior Editor, AI & Product
Cecilia leads the Swarmz editorial desk. She has spent a decade turning complex AI and product topics into writing people actually finish, and she owns the blog's quality bar.


