Why AI Models Repeat Themselves, and How to Stop It
Repetition has three separate causes, and only one of them is fixed by temperature. How to tell a sampling loop from context echo from your own prompt.
Why AI Models Repeat Themselves, and How to Stop It
If you are wondering why AI repeats itself, the honest answer is that there are three separate causes and they get confused with each other constantly. One is a sampling problem, one is a context problem, and one is a prompt problem. Turning the temperature up only helps with the first, which is why that advice so often does nothing.
Work out which one you have before you change a setting.
Cause one: the sampling loop
A language model picks each token from a probability distribution. When that distribution is very peaked, the same token wins every time, and if the resulting text feeds back into a state that produces the same peaked distribution again, you get a loop. The classic symptom is a sentence, clause or phrase repeating verbatim and endlessly, sometimes until the output limit cuts it off.
This one is mechanical and it has mechanical fixes:
Control | What it does | When to reach for it |
|---|---|---|
Temperature | Flattens or sharpens the whole distribution | Output is locked in a verbatim loop at a low setting. Raise it in small steps |
Top-p | Truncates the distribution to the smallest set of likely tokens | A very low top-p can cause the loop. Widening it gives the model somewhere else to go |
Frequency penalty | Reduces the probability of tokens already used, in proportion to their count | The right tool when a specific word or phrase keeps coming back |
Presence penalty | Reduces the probability of any token already used at all | Useful when you want topical variety rather than vocabulary variety |
If you are not sure what those two sampling numbers actually do, temperature and top-p are worth understanding properly before you start turning them, because changing both at once makes the result impossible to attribute.
Two practical notes. Reasoning models loop far less often than they used to, so if you are seeing a hard verbatim loop on a current model it is worth checking whether the request is being served by a much smaller or older model than you think. And a frequency penalty set too high produces its own failure mode, where the model starts avoiding words it needs, including technical terms it should reuse.
Cause two: context echo
This one looks like repetition but is not a sampling failure at all. In a long chat, the model repeats a paragraph it wrote earlier, or restates a conclusion it already gave, because that text is in its context and the most probable continuation of a conversation containing a summary is another summary.
The tell is that the repeated content is something from earlier in the same conversation, reworded rather than duplicated character for character. Temperature will not touch this. The fixes are structural:
Start a new conversation. The single most effective fix, and the most under-used.
Remove the earlier draft from the context rather than asking for a new one below it.
Ask for a diff rather than a rewrite, so the model is not re-emitting the whole passage.
Keep the running context shorter. Degradation over a long context is a real effect, not a superstition.
Long conversations degrade in other ways too, and repetition is often the first visible symptom rather than the only one. Why a chatbot starts forgetting earlier instructions covers the same underlying pressure from the other direction.
Cause three: instruction restatement
The third kind is the model repeating your own instructions back at you, or repeating a framing you supplied. You ask for a concise summary and each paragraph opens by announcing that it is a concise summary. This is not degenerate output. It is the model treating your instruction as content because you put it in a position where content goes.
The fix is prompt structure. Separate the instruction from the material clearly, put the instruction in a system prompt where one is available, and stop asking for the same property in three different sentences. Three requests for brevity in one prompt make brevity the topic.
Anthropic's prompt engineering documentation makes the same point about separation, and the practical technique for keeping the two apart is in using XML tags in AI prompts.
Diagnosing it in thirty seconds
Is the repeated text word for word identical, and does it continue until the output ends? That is a sampling loop. Adjust top-p first, then temperature.
Is the repeated text something from earlier in the same conversation, reworded? That is context echo. Start a fresh conversation.
Is the model restating your instruction or your framing? That is prompt structure. Move the instruction out of the content.
One more thing worth ruling out first. If the output stops mid-sentence as well as repeating, you may be hitting an output limit rather than a repetition problem, which is a different diagnosis covered in why AI responses get cut off.
All three causes sit on top of the same token-by-token generation process, and a working mental model of that process makes the distinction between them obvious rather than mysterious. How AI models work is the place to start if the sampling vocabulary above was unfamiliar.
Frequently asked questions
Why does raising temperature sometimes make repetition worse?
Because it is the wrong lever for your cause. If the repetition is context echo, a higher temperature gives you a more randomly worded version of the same repeated idea, which reads as both repetitive and sloppier. It genuinely feels worse, and it is.
Should I always set a frequency penalty?
No. Leave it at zero by default and reach for it when you have an identified phrase problem. Applied globally it degrades any output that legitimately needs to reuse a term, which includes most technical writing, legal text and code.
Does repetition mean the model is too small for the task?
Sometimes. Smaller models loop more, especially at low temperature and long output lengths. If the same prompt at the same settings loops on a small model and does not on a larger one, that is a real signal rather than a coincidence.
Why does it repeat more at the end of a long response?
The further into a generation you are, the more of the context is the model's own output. That makes the next token increasingly conditioned on text it wrote rather than text you wrote, which amplifies whatever pattern it has settled into. Shorter outputs, requested in sections, avoid most of it.
How did this land?
About the author

Staff Engineer, Platform
Carlo works on the platform that turns prompts into running apps. He writes the engineering deep dives and the changelog notes worth reading.


