Context Rot

Intermediate

How length, position and distractors can affect the use of context.

Last updated: Sep 13, 2026

What context rot means

A large context window describes how much input a model can accept. It does not guarantee reliable use of every detail. Depending on the model and task, extra documents may help, distract or make relevant information harder to use. This is empirical behavior, not a rule that quality must decline with every token.

Lost in the Middle: a concrete experiment

Liu et al. (2023, revised 2024) tested NaturalQuestions-Open questions. In the multi-document task, one Wikipedia passage contained the answer and other passages served as distractors. They varied the answer passage position and used greedy decoding. Accuracy checked whether an accepted answer appeared in the output.

Several models tested then did better when the answer passage was near the start or end than in the middle. These curves measure task accuracy, not attention weights. Other tasks and models showed different patterns.

Original values from Table 1: control conditions, not a position curve
ModelClosed-bookOracle
LongChat-13B (16K)35%83.4%
MPT-30B-Instruct31.5%81.9%
GPT-3.5-Turbo (0613)56.1%88.3%
Claude-1.348.3%76.1%

Historical measurements for these model versions and task. They do not predict current models or your own documents.

Liu et al., §2, Table 1, Figure 5

Test length and distraction separately

Chroma’s Context Rot study (2025) compared 18 models while varying input length and distractors. Results depend on model and setup. A synthetic retrieval task does not automatically represent contradictory sources or long conversations.

Chroma: Context Rot

Compute cost is not task accuracy

Full score matrix per head
1,000,000
n²
Causally allowed entries including diagonal
500,500
n(n+1)/2

These count dense self-attention scores across a full sequence, not measured runtime. FlashAttention need not materialize the full matrix in main memory; a decode step with KV cache has different costs. Learned attention is not uniform: a token can receive high weight even in a long context.

How to test your own workflow

  1. Define questions, reference answers and relevant sources. Fix model version, prompt and decoding settings.
  2. Compare the answer source alone, the same source with distractors, and different positions at equal total length.
  3. Measure answer correctness and source support separately. Repeat stochastic runs and record variation, cost and latency.
  4. Test retrieval, structured notes or compaction on the same cases. Check whether important qualifications are lost.