Overview
The industry sold long context as a solved problem. Million-token windows, near-perfect Needle-in-a-Haystack scores, and the promise that you could dump a corpus into the prompt and let the model "just use it."
That promise does not survive contact with either controlled evaluation or real agentic search.
Two independent lines of research now make the same claim from different directions. Chroma's July 2025 technical report, Context Rot: How Increasing Input Tokens Impacts LLM Performance, showed that even simple tasks degrade as input length grows — and that lexical NIAH is a misleading benchmark. Xia, Wang, Huang and Liu (arXiv:2606.29718, August 2026) then showed what that degradation looks like in long-horizon search agents: premature termination. Under extensive context, models give up or emit uncertain incorrect answers long before they exhaust the context window.
If you are stuffing RAG dumps, conversation histories, or multi-hop search traces into a large window and assuming the model will use them, you are designing against the evidence.
The Study
Hong, Troynikov and Huber (Chroma, July 2025) evaluated 18 LLMs — including GPT-4.1, Claude 4, Gemini 2.5 and Qwen3 — on tasks deliberately simpler than production RAG. They extended Needle-in-a-Haystack beyond lexical matching, added LongMemEval conversational QA, and a synthetic repeated-words task designed to isolate length as the only variable.
Xia et al. (2026) then moved the question into the setting that actually matters for agents: deep search. They studied four flagship open models with 200K–256K windows — GLM-4.7, GLM-5.0, Qwen3.5-397B-A17B and MiniMax-M2.5 — across BrowseComp, BrowseComp-Plus and xbench-DeepSearch. Closed models were excluded because they encrypt reasoning traces, which made the terminal-state analysis infeasible. A GPT-OSS-120B judge classified outcomes with 98.7% agreement against 300 human-annotated trajectories.
Together, the two papers answer different halves of the same question: does more context help? Chroma answers for single-turn use. Xia et al. answer for multi-turn agents.
What They Found
1. Perfect NIAH scores are not evidence of uniform context use
Standard Needle-in-a-Haystack is a lexical retrieval task: a known sentence is planted in unrelated text and the model is asked to retrieve it. Models score near-perfectly. Chroma's point is that this is the wrong test. When the needle is a semantic match rather than a verbatim one, and when the haystack is varied (Paul Graham essays vs arXiv papers), performance degrades as length grows — non-uniformly, and in ways that depend on needle–haystack similarity and distractor difficulty.
The implication is uncomfortable: the benchmark the industry used to declare long context "solved" was measuring the easiest version of the problem.
2. Degradation appears even on trivial tasks
Chroma's repeated-words task asks models to reproduce a sequence of repeated tokens. There is no reasoning, no retrieval ambiguity, no tool use. Performance still falls as input length increases, across Claude Sonnet 4, GPT-4.1, Qwen3-32B and Gemini 2.5 Flash. If a model cannot reliably copy under length pressure, it will not reliably reason over a 200K agent trace.
3. In agents, the failure is not "running out of context" — it is giving up
Xia et al. taxonomise terminal states:
| State | Definition | |---|---| | Give up | The agent states it cannot solve the problem and does not give a clear answer | | Uncertain answer | A clear answer, but the reasoning explicitly flags unresolved uncertainty | | Confident answer | A clear answer, with reasoning that claims all criteria are met | | No answer | The agent hits the context limit or turn budget |
On BrowseComp, give-up rates are not a rounding error. GLM-4.7 gave up on 38.6% of trajectories. Qwen3.5-397B-A17B gave up on 19.8%. MiniMax-M2.5 gave up on only 5.2% — but produced uncertain incorrect answers on 48.2%. Hitting the context limit ("No answer") was essentially 0% for most model–dataset pairs.
The models are not dying because the window is full. They are aborting, hedging, or confidently guessing while there is still room to search.
4. Premature termination rises with context length, even when difficulty is held constant
This is the result that should change system design. Xia et al. control for query difficulty and still find that premature termination rate is positively correlated with trajectory length. Longer traces do not merely contain harder problems. Length itself induces give-up and uncertain-incorrect behaviour.
Accuracy drops sharply as trajectory length increases. Confident incorrect answers dominate early; give-up and uncertain-incorrect dominate later. That is a different failure curve from "lost in the middle," and it will not be fixed by putting the answer at the start of the prompt.
5. Context management is test-time scaling, not a free lunch
The authors evaluate seven context-management methods across three families: compaction (summarise history), trimming (drop older turns), and isolation (sub-agent calls that keep the parent context clean). These methods reduce premature termination and improve accuracy — and they also increase cost and unfinished trajectories. They are buying more exploration with more inference.
Method choice is model-dependent:
- Strong agentic backbones: context isolation via sub-agent calls outperforms other methods.
- Weaker agentic backbones: compaction + trimming is the better performance/cost balance.
Parallel sampling, with a behaviour-aware filter that drops give-up and uncertain-answer trajectories before aggregation, adds 2.6–4.9% across three aggregation methods — without changing the ReAct loop.
Why This Happens
Three mechanisms, stacked:
-
Softmax and attention are not length-invariant. This is the same structural story as attention sinks and lost-in-the-middle: as sequence length grows, attention mass concentrates, middle content is starved, and the model’s effective working memory is much smaller than the advertised window. Chroma’s non-uniform NIAH results are the single-turn expression of that fact.
-
Training did not prepare models for 100-turn search traces. Pre-training documents are not agent trajectories. Instruction-tuning examples are not 80-page concatenations of search snippets, tool JSON, and self-talk. When the distribution shifts far enough, the model’s learned stopping rules fire early: “I cannot find this,” “I should be honest,” “good enough.”
-
Honesty training collides with exploration. Give-up is not always a bug in isolation — calibrated refusal is useful. In deep search it becomes a failure mode, because the honest-looking abort happens before the search is exhausted. Xia et al.’s judge found give-up and uncertain-answer states highly correlated with incorrect final answers. The model is not wisely abstaining. It is stopping the search.
Concrete Recommendations
1. Stop treating context window size as a capability
A 1M window is a maximum, not a working memory. Budget a much smaller effective context — empirically, well below the point where your model’s premature-termination curve steepens — and design retrieval, memory and sub-agents around that budget.
2. Do not dump RAG results in relevance order
Chroma shows semantic similarity between needle and haystack changes degradation. Xia et al. show length itself induces abort. Combined with Liu et al. (2023) on lost-in-the-middle: retrieve fewer, better chunks; put the highest-value evidence first and last; never pad.
3. Instrument terminal states, not just accuracy
If you only log pass/fail, you cannot tell a retrieval miss from a give-up. Classify every agent trajectory as give-up / uncertain / confident / truncated. Give-up rate as a function of token count is the metric that tells you whether you have a context-rot problem.
4. Match context strategy to model strength
Do not copy a “sub-agent isolation” architecture onto a weak backbone and expect Xia et al.’s result. Isolation wins for strong agentic models. Compaction plus trimming is the cheaper default for weaker ones. Measure both accuracy and unfinished-trajectory rate; the paper is explicit that mitigations buy exploration with cost.
5. Filter parallel samples by behaviour, not by vote alone
Majority vote over traces that include give-ups will launder aborts into the answer. Drop give-up and uncertain-answer trajectories before aggregation. The 2.6–4.9% gain is modest, but it is free of architectural change.
6. Keep search traces out of the parent context
For deep research / multi-hop search products, the parent agent should see a compressed brief, not the raw page dump. Isolation is the method Xia et al. recommend for strong models precisely because it prevents the parent from accumulating the length that triggers premature termination.
How This Extends Prior LLM Wisdom
This is not a restatement of lost-in-the-middle or attention sinks. Those papers explain where information is ignored inside a single prompt. Context rot explains what happens as you keep adding tokens anyway — including in multi-turn agents that never hit the window limit. Premature termination is the operational name for that failure.
If you already restructured prompts around U-shaped attention, the next design move is to stop growing the prompt.
Follow-On Research
- Hong, Troynikov & Huber (2025). Context Rot: How Increasing Input Tokens Impacts LLM Performance. Chroma technical report.
- Xia, Wang, Huang & Liu (2026). Diagnosing and Mitigating Context Rot in Long-horizon Search. arXiv:2606.29718.
- Liu et al. (2023). Lost in the Middle: How Language Models Use Long Contexts. arXiv:2307.03172.
- Hsieh et al. (2024). RULER: What’s the Real Context Size of Your Long-Context Language Models? arXiv:2404.06654.
Recommended Reading
- Chroma Context Rot report
- arXiv:2606.29718 and code
- LLM Wisdom: The Lost in the Middle Problem
- LLM Wisdom: Attention Sink Phenomenon

