What context rot means
A large context window describes how much input a model can accept. It does not guarantee reliable use of every detail. Depending on the model and task, extra documents may help, distract or make relevant information harder to use. This is empirical behavior, not a rule that quality must decline with every token.
Lost in the Middle: a concrete experiment
Liu et al. (2023, revised 2024) tested NaturalQuestions-Open questions. In the multi-document task, one Wikipedia passage contained the answer and other passages served as distractors. They varied the answer passage position and used greedy decoding. Accuracy checked whether an accepted answer appeared in the output.
Several models tested then did better when the answer passage was near the start or end than in the middle. These curves measure task accuracy, not attention weights. Other tasks and models showed different patterns.
| Model | Closed-book | Oracle |
|---|---|---|
| LongChat-13B (16K) | 35% | 83.4% |
| MPT-30B-Instruct | 31.5% | 81.9% |
| GPT-3.5-Turbo (0613) | 56.1% | 88.3% |
| Claude-1.3 | 48.3% | 76.1% |
Historical measurements for these model versions and task. They do not predict current models or your own documents.
Liu et al., §2, Table 1, Figure 5Test length and distraction separately
Chroma’s Context Rot study (2025) compared 18 models while varying input length and distractors. Results depend on model and setup. A synthetic retrieval task does not automatically represent contradictory sources or long conversations.
Chroma: Context RotCompute cost is not task accuracy
- Full score matrix per head
- 1,000,000
- n²
- Causally allowed entries including diagonal
- 500,500
- n(n+1)/2
These count dense self-attention scores across a full sequence, not measured runtime. FlashAttention need not materialize the full matrix in main memory; a decode step with KV cache has different costs. Learned attention is not uniform: a token can receive high weight even in a long context.
How to test your own workflow
- Define questions, reference answers and relevant sources. Fix model version, prompt and decoding settings.
- Compare the answer source alone, the same source with distractors, and different positions at equal total length.
- Measure answer correctness and source support separately. Repeat stochastic runs and record variation, cost and latency.
- Test retrieval, structured notes or compaction on the same cases. Check whether important qualifications are lost.