The Study
Liu, Nelson, Lin, Sewon, Liang, and Manning (2023) published Lost in the Middle: How Language Models Use Long Contexts (arXiv:2307.03172), examining how 13 commercially-deployed and open-source LLMs use information positioned at different points within a context window.
The findings are striking and largely underreported in practitioner communities.
What They Found
Using multi-document question answering tasks, the researchers varied the position of the relevant document within a set of 10–30 distractor documents. Performance was measured as exact-match accuracy on factual questions.
The U-shaped degradation curve: Models consistently performed best when relevant information appeared at the very beginning or very end of the context. Performance dropped sharply — in some models by 20+ percentage points — when relevant information was placed in the middle.
This was observed across:
- GPT-3.5-Turbo (16K)
- Claude 1.3 (100K)
- LongChat-13B
- MPT-30B-Instruct
- And 9 other models
Longer context = worse middle performance: Counterintuitively, models with larger context windows showed more pronounced U-curves. More room to attend = more opportunity for the middle to be ignored.
Why This Happens
The authors hypothesise several interacting causes:
-
Training data distribution bias: Most pre-training text has relevant information near the start (news articles, abstracts, introductions). Models inherit this positional bias.
-
Attention decay: As sequence length increases, attention scores for middle positions tend to be smaller in magnitude, even with positional embeddings designed to handle long sequences.
-
Instruction following vs. retrieval: Models fine-tuned for instruction following may have learned shortcuts (answer from recent context, or from strong priors) that bypass mid-sequence retrieval.
Concrete Recommendations
1. Put the most important content first or last — never in the middle
For RAG applications: place the most relevant retrieved chunk at position 0 or as the final chunk before the question. Do not let relevance scoring alone determine ordering — consider also positioning.
2. For few-shot prompts, the last example matters most
Recency bias means your most representative example should be the one immediately before the task. If you have one strong example and several mediocre ones, put the strong one last.
3. Compress, don't pad
If you're filling context with retrieved documents, fewer high-quality chunks outperform many lower-quality ones. Padding context with marginally relevant material actively degrades performance by pushing relevant content toward the middle.
4. System prompts are not read equally throughout
Long system prompts suffer from the same effect. Critical instructions (output format, constraints, persona) should appear at the beginning and be briefly restated at the end. Instructions buried mid-prompt are frequently violated.
5. Test with position-scrambled evaluation
When building RAG pipelines, benchmark performance not just on overall accuracy but on accuracy conditioned on retrieval rank position. If your best-performing retrieval position is always rank-1, you have a positioning problem, not a retrieval problem.
Follow-On Research
An et al. (2024) extended this work in Make Your LLM Fully Utilize the Context (arXiv:2404.16811), showing that fine-tuning on position-diverse datasets partially mitigates the U-curve but does not eliminate it. No current model is immune.
Recommended Reading
- Liu et al. (2023). Lost in the Middle: How Language Models Use Long Contexts. arXiv:2307.03172
- An et al. (2024). Make Your LLM Fully Utilize the Context. arXiv:2404.16811
- Press et al. (2022). Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation. arXiv:2108.12409
