Lost in the Middle: Why AI Misses Facts Mid-Prompt
Language models often use the start and end of a long prompt better than the middle. Here is the research, a ten-minute test you can run yourself, and the fixes that work.
The lost in the middle problem is the finding that language models tend to use information at the start and end of a long prompt better than information buried in the middle. It was documented in the 2023 paper Lost in the Middle: How Language Models Use Long Contexts, and it changes a practical decision: where you place a fact in a prompt can matter as much as whether the fact fits.
If you paste a forty-page policy into a chat and ask about clause 17, that clause sits right in the zone the research flagged. This post covers what was measured, a test you can run on your own model in ten minutes, and what to do with the result.
What the research measured
The authors tested models on two kinds of task: multi-document question answering and key-value retrieval. In both, they moved the position of the relevant information inside the prompt and watched accuracy.
The result was a U-shaped curve. Accuracy was often highest when the answer sat at the beginning or the end of the context, and it dropped, sometimes sharply, when the answer sat in the middle. Adding more documents made the effect more visible.
Two honest caveats. The paper tested the models available in 2023, so it tells you nothing certain about the one you use today. And the causes are still debated, so be wary of anyone who gives you one tidy explanation. What the paper gives you is a testable behavior, not a law.
A ten-minute position test you can run yourself
You do not need a benchmark suite. You need one long document and one fact the model cannot possibly know. The invented fact matters: if you use a real fact, the model may answer from training data and the test measures nothing.
Take a long text, such as a concatenation of 15 to 20 pages of your own documentation. Aim for a length you actually use in practice.
Write one invented sentence, for example: The staff entrance code word is teal-harbor-47.
Build three versions of the prompt. Put the invented sentence near the top, near the exact middle, and near the bottom.
After the document, ask: What is the staff entrance code word? Answer only from the text above.
Run each version five times in fresh chats. Count the correct answers per position.
[document part 1]
...
The staff entrance code word is teal-harbor-47. <- move this line
...
[document part 2]
Question: What is the staff entrance code word?
Answer only from the text above. If it is not there, say so.Record the tally in a small table. Five runs per position is a rough probe, not a statistic, but a gap like 5 of 5 at the top against 1 of 5 in the middle tells you something real.
Position of the fact | Runs | Correct answers |
|---|---|---|
Top of the document | 5 | fill in |
Middle of the document | 5 | fill in |
Bottom of the document | 5 | fill in |
If every position scores the same, good: your model handles this document length well for this kind of lookup. Try again with a longer document or a subtler question, because a lookup of one distinctive sentence is the easy case.
What to do if the middle loses
The fixes are mostly about not making the model hunt.
Put the instruction and the question last. Repeat the question after the documents, not only before them.
Send less, better material. Retrieval that returns five relevant chunks beats a dump of fifty. The tradeoffs are covered in RAG vs fine-tuning vs long context.
Order the evidence. If you rerank retrieved chunks, place the strongest at the start or the end, not in the center.
Split the job. Summarize or extract from each section in a separate call, then combine the short results.
Quote first. Ask the model to copy the relevant sentences before answering. Anthropic's prompting guide suggests quoting the relevant parts of long documents first so the model focuses on the relevant content and ignores the rest.
Longer prompts also cost more per call, which gives you a second reason to trim. See why long context costs more for how that bill builds.
What lost in the middle does not mean
It does not mean long context windows are useless. A large window lets a document fit at all, and for many questions the model does fine. Does a bigger context window mean better answers takes that question apart properly.
It is also a different failure from running out of room. When the prompt exceeds the limit, text gets cut, which is covered in what happens when AI runs out of context window. Lost in the middle happens while everything still fits.
For a broader map of how prompts, windows and retrieval fit together, the how AI models work guide is the place to start, and what is context engineering covers the habit of curating what goes into the window.
FAQ
Does every AI model have the lost in the middle problem?
Not to the same degree, and not on every task. The original paper found it across the models it tested in 2023. The only reliable answer for your model is a position test like the one above.
Is lost in the middle the same as a context window limit?
No. A context window limit is a hard cap on how much text the model can read. Lost in the middle is about uneven attention to text that fits inside the cap.
Where should I put the most important instruction in a long prompt?
At the end, directly before the model starts answering, and repeat the key constraint at the start if the prompt is long. Then test it, since the right spot varies by model.
Will RAG fix the problem?
It helps because it shortens the prompt, so there is less middle to get lost in. But a retrieval step that returns too many chunks recreates the same long prompt.
How did this land?
About the author

Senior Editor, AI & Product
Cecilia leads the Swarmz editorial desk. She has spent a decade turning complex AI and product topics into writing people actually finish, and she owns the blog's quality bar.


