Why AI Models Get Slower in Long Conversations
There is no session. The model re-reads the entire conversation every turn, and attention cost grows with the square of the length. Prefill is what you are waiting for.
Why AI Models Get Slower in Long Conversations
AI models get slower in long conversations because a model has no memory between turns. Every time you send a message, the entire conversation so far is fed back through the model as input, and the cost of processing that input grows faster than the conversation does. A chat that felt instant at message five can take several seconds to start responding at message fifty, even though your last message was three words long. The slowdown is in the reading, not the writing.
The part that surprises people: there is no session
It is natural to picture a chat as a conversation the model is following, adding each new message to something it already knows. That is not what happens. The model is stateless. The chat interface keeps a transcript and resends all of it, every single turn, along with the system instructions and anything else stuffed into the context window. Turn fifty is not one short message, it is fifty messages plus their replies, submitted as one large block of text.
So the growth is quadratic in a specific sense. Each turn is longer than the last, and you are paying the processing cost of that growing block on every single turn. Over a long conversation the total work done is proportional to the square of the conversation length, which is why the slowdown feels sudden rather than gradual: it is fine, fine, fine, then noticeably worse.
Prefill versus decode
Splitting the wait into its two phases, prefill and decode, explains everything you observe.
Phase | What happens | How it scales |
|---|---|---|
Prefill | The model reads the entire conversation and builds internal state | Grows with total conversation length, dominates long chats |
Decode | The model generates its reply, one token at a time | Grows with reply length only, roughly flat per token |
The delay before the first word appears is prefill. The speed at which words then stream out is decode. In a long conversation, decode speed barely changes while prefill balloons, which is exactly the experience of waiting a long time and then getting a fast, normal-looking answer. If you want the precise metric for the first half, it is time to first token, and it is the number that degrades.
Attention makes prefill worse than a simple word count would suggest. In the transformer architecture introduced in 2017, every token attends to every other token, so the attention work scales with the square of the input length rather than linearly with it. Modern implementations optimise this heavily, but the shape of the curve survives the optimisations.
Why it also gets more expensive
The same mechanism explains the bill. You are charged for input tokens on every turn, and the input is the whole conversation. A fiftieth turn can cost more than the first forty combined, which is the same reason long context costs more in any application, not just chat. Cost and latency here are two readings of one underlying fact.
What actually fixes it
Start a new conversation. The cheapest fix and the most underused. Carry forward a short summary of what matters rather than the full history, which is also what you would do if the conversation hit the window limit outright.
Use prompt caching if you are building on an API. When a long prefix stays identical between calls, providers can cache the processed state and skip most of the prefill. Prompt caching is the single biggest latency win available for long system prompts and stable document context, and it is often left switched off.
Prune rather than append. Drop resolved tangents, failed attempts and pasted material you no longer need. Most long conversations are mostly dead weight by the halfway point.
Do not paste large documents into chat repeatedly. A document pasted at turn three is re-read at every turn thereafter. Retrieval, or a single summarisation pass, replaces fifty re-reads with one.
What does not fix it is a bigger context window. A larger window raises the ceiling on how much you can include; it does not make processing that material faster or cheaper, and the answer to whether a bigger context window means better answers is largely no. Filling a big window is what causes the problem this post describes.
FAQ
Does the model remember my earlier messages?
Not on its own. The application resends the transcript each turn, which is what creates the impression of memory and also what creates the slowdown. Some products add a separate long-term memory feature, but that is a retrieval system layered on top, not the model remembering.
Why does the reply start slowly but then stream quickly?
The slow part is prefill, where the model reads the whole conversation. The fast part is decode, which is unaffected by conversation length. That split is the clearest everyday evidence for the mechanism.
Is a longer conversation also less accurate?
It can be. Long inputs make it easier for a model to lose track of an instruction given early on, which is a separate problem from speed but has the same fix: start fresh and carry a summary.
Does this affect AI coding agents too?
Yes, and more visibly, because agents accumulate file contents and command output fast. The practical symptom is an agent that slows down and starts contradicting itself, and the recovery is the same checkpoint-and-restart move: summarise the state, start a fresh session.
How did this land?
About the author

Senior Editor, AI & Product
Cecilia leads the Swarmz editorial desk. She has spent a decade turning complex AI and product topics into writing people actually finish, and she owns the blog's quality bar.


