In short
When an LLM’s context grows longer, its performance drops — and this drop is measurable, reproducible, and mathematically modelable. Liu et al. (2024) documented a U-shaped positional bias: the model handles information at the beginning or end of the context better, and information in the middle poorly. Du et al. (2025) isolated raw length as an independent causal variable: even with perfect retrieval, more tokens degrade performance. Roy et al. (2026) formalize the dynamics: exponential decay for standard Transformers, polynomial decay for decomposed architectures.
In short: imagine an assistant to whom you read a file aloud. At the beginning, they listen attentively. At the end, they better retain what they have just heard. But in between, their attention scatters — and the longer the file, the wider the “forgotten middle” stretches. This phenomenon is not an impression: it is measured, modeled, and shapes the design of any serious RAG system.
The problem: the longer the context, the lower the quality
LLMs (large language models) operate within a context window — the maximum amount of text they can process in a single pass. This window has expanded considerably in recent years: from 4,096 tokens in 2023 to 200,000 tokens or more for some current models.
This expansion fueled an implicit assumption in RAG (Retrieval-Augmented Generation) systems: if the model has all the relevant information in its context, performance should remain stable. Recent work invalidates this assumption.
Hong, Troynikov and Huber (Chroma Research, 2025) tested 18 frontier models — GPT-4.1, Claude Opus 4, Claude Sonnet 4, Gemini 2.5, Qwen3, and others — across several long-context retrieval benchmarks. The result was uniform: every tested model exhibits context rot at every increment of input length. The term context rot designates this progressive degradation in response quality as context length increases.
The magnitudes documented by Chroma: precision drops of 20 to 50% between 10,000 and 100,000 tokens on NIAH (Needle in a Haystack) tasks — retrieving a target sentence from a long document. A model with a declared window of 200,000 tokens can show significant degradation at 50,000 tokens — just one quarter of its theoretical capacity.
In short: the window advertised by a model (200K tokens) is a technical ceiling, not a zone of constant quality. It is the equivalent of a car sold to drive at 250 km/h: it technically can, but stability, braking, and consumption degrade well before. Designing a RAG by assuming the model uniformly handles 100K tokens is to ignore the measured degradation.
Lost in the Middle — the U-shaped curve
Liu et al. (2024) were among the first to rigorously document the positional bias of LLMs. Their study, published in the Transactions of the Association for Computational Linguistics, covers two tasks: multi-document question answering and key-value retrieval.
The central result is a U-shaped performance curve as a function of the position of the relevant information in the context:
- Beginning of context (primacy bias): maximum performance.
- End of context (recency bias): maximum performance.
- Middle of context: minimum performance, even for models explicitly trained on long contexts.
The numbers are precise. On the multi-document QA task, performance drops by more than 30% when the answer document moves from position 1 to position 10 in a 20-document context. GPT-3.5-Turbo, with the information placed in the middle, falls below its closed-book score — that is, below what it would achieve with no documents provided at all. On key-value retrieval, some models achieve near-perfect accuracy at extreme positions and drop below 40% in the middle.
This phenomenon is not trivial: it means that context can actively hurt performance by introducing positional noise that exceeds the value of the information provided.
Probable mechanism
Liu et al. formulate a mechanistic hypothesis: instruction-tuning data typically places the instruction at the beginning of the context. The model structurally learns to prioritize the beginning — and, by attention symmetry, the end. The middle becomes a zone of low attention.
This hypothesis is not experimentally confirmed in the paper. The authors acknowledge that more recent positional encoding techniques (RoPE, ALiBi) may alter this profile, and that their results cover 2023-era models.
In short: placing a critical instruction in the middle of a long prompt is like saying the essential to an audience that is bored in the third act of a play. For a directive to be followed, it must be at the beginning or end of the context — not buried in the mass.
Contextual degradation follows a mathematical law — the λ-RLM model
Roy et al. (2026) provide the first rigorous mathematical formalization of context rot. Their paper “The Y-Combinator for LLMs” (arXiv:2603.20105) establishes two decay models, depending on the system architecture.
Exponential decay: the standard Transformer case
For a standard Transformer operating in direct mode (Direct M), accuracy follows exponential decay:
A_direct(n) = Θ(ρ^(n/K)) → 0
Where n is the context length, K the model’s native window, and ρ ∈ (0,1) the decay rate. In plain terms: when n greatly exceeds K, accuracy converges to zero. The degradation accelerates — each doubling of context beyond the native window multiplies the loss by a constant factor.
Polynomial decay: the λ-RLM case
The λ-RLM (Lambda Recursive Language Model) architecture replaces the open inference loop with a typed functional runtime, grounded in λ-calculus. The principle: decompose the long problem into subproblems bounded within the native window, process each subproblem independently, then aggregate. The library of predefined combinators — SPLIT, MAP, FILTER, REDUCE — guarantees termination and cost bounds.
In this case, decay follows a power law:
A_λ-RLM(n) = Ω(n^(−c))
Where c = −log_{k*}(A(τ*) · A_⊕), with k* the optimal branching factor, A(τ*) the accuracy on subproblems, and A_⊕ the aggregation accuracy. A power law decays much more slowly than an exponential: at 10× the native window, an exponential may approach zero while the power law retains a significant fraction of initial accuracy.
In the ideal case — a perfectly decomposable problem with zero aggregation errors — accuracy remains constant regardless of length n.
Empirical results
Across 4 long-context reasoning tasks and 9 base models, λ-RLM outperforms standard RLMs in 29 out of 36 comparisons. The largest gains concern “weak” models: average improvement of up to +21.9 accuracy points. An 8B model scaffolded with λ-RLM exceeds the accuracy of a standard 405B model on long-context tasks. At equivalent accuracy (8B vs 70B standard), λ-RLM is 3.1× faster.
Model limits
The formal guarantees apply to the defined combinator library — extension to arbitrary combinators is not covered. The model assumes that decomposition and aggregation errors are independent: an assumption that can be violated when subproblems share semantic dependencies. Evaluations cover 4 specific tasks; generalization to open-ended code generation or multimodal tasks remains to be validated.
Length ≠ quality: when more tokens degrade performance
Du et al. (2025) pose a more radical question: is raw length itself a cause of degradation, independent of content?
Their protocol isolates the length variable. The experimental structure is: [Evidence] [Distractor tokens] [Question], with the evidence placed consecutively. Five models are tested: Llama-3.1-8B-Instruct, Mistral-v0.3-7B-Instruct, GPT-4o, Claude-3.7-Sonnet, Gemini-2.0. Tasks cover mathematics, question answering and coding.
Raw length is causal
The central result: input length alone degrades performance, independently of retrieval quality. Llama-3.1-8B-Instruct drops by 24.2% on the MMLU benchmark extended to 30,000 tokens, despite a perfect retrieval rate of 97% — the right passage is present, the model locates it, but its performance degrades nonetheless.
The critical experiment: replacing natural-language distractor tokens with whitespace tokens, semantically neutral. The degradation persists:
- Llama: −48% on VarSum at 30,000 tokens.
- Mistral: −30% on GSM8K at 30,000 tokens.
This result rules out semantic distraction as the sole cause. It is not that the model is confused by confusing content — it is the raw quantity of tokens that degrades attention.
Observed degradation ranges from 13.9% to 85% across models and tasks, at lengths that remain well below declared context windows.
Implications for RAG systems
These results explain why RAG performance saturates — or even regresses — when more documents are added to the context. The strategy “more context = more information = better answer” is incorrect beyond a certain threshold. They also corroborate observations on long chains of thought (long chain-of-thought) that can hurt performance on simple tasks.
Du et al. propose a partial mitigation: recitation prompting, which consists of asking the model to paraphrase the retrieved evidence before answering. Observed gain on GPT-4o: +4% on the RULER benchmark — real, but insufficient to address the structural cause.
In short: length alone tires the model, like an orator who must keep concentration over three hours of speech. Even if the key sentence is right there, just under their eyes, the processing quality drops. Operational solution: do not load 30,000 tokens “just in case” — provide the 5,000 tokens that are truly useful.
Architectural solutions: decompose to preserve
All four sources converge on an operational conclusion: the solution is not to increase the context window, but to decompose the problem so that each call stays within the model’s optimal performance zone.
Focused prompts vs. full prompts
Chroma Research (2025) documents the most immediately actionable effect: focused prompts outperform full prompts by 30 to 60% on retrieval tasks. Providing the model with only the relevant passage rather than the entire document drastically improves accuracy — even if the context window could absorb the full document.
Another counter-intuitive result from Chroma: models perform better on randomly shuffled haystacks than on logically coherent documents. The authors’ hypothesis: thematic coherence increases semantic interference between nearby passages.
Functional decomposition (λ-RLM)
The λ-RLM approach from Roy et al. formalizes this principle at the architectural scale. Decomposition is not ad hoc — it is structured by typed combinators with formal termination guarantees. Each leaf of the execution graph receives a subproblem calibrated to the native window of the base model.
The trade-off: this architecture requires an external runtime and imposes decomposability of the problem. Tasks that do not fragment cleanly — continuous narrative reasoning, long-form stylistic coherence — do not benefit from the same guarantees.
What current architectures do not solve
Models with extended windows (128K, 200K tokens) reduce the problem without eliminating it. Chroma (2025) shows that a 200K-window model can degrade significantly at 50K tokens. Attention remains a quadratic mechanism in complexity: in a sequence of 100,000 tokens, each token shares its attention budget across 99,999 neighbors. Extending the window pushes the degradation threshold back — it does not eliminate the curve.
Decision matrix — which strategy for which context
| Application context | Recommendation | Why |
|---|---|---|
| Factual Q&A, corpus < 50K tokens | RAG + focused prompts (top-5 chunks) | Stays in the model’s optimal zone, +30 to 60% precision vs. full prompt (Chroma 2025). |
| Single long document, sequential reading | Native long context up to 50% of the declared window | Beyond, Chroma degradation ≥ 20%. Better to section. |
| Multi-step reasoning over large corpus | λ-RLM decomposition or agent + sub-tasks | Polynomial decay instead of exponential, documented gains up to +21.9 points. |
| Critical information to place in the prompt | Beginning or end of context (never middle) | U-shaped positional bias, loss > 30% in the middle (Liu 2024). |
| Strict latency, non-negotiable quality | Smaller model + short context > large model + long context | An 8B λ-RLM beats a 405B standard on long tasks at 3.1× the speed. |
What to remember
- Contextual degradation is measurable and reproducible: Liu et al. (2024), Du et al. (2025) and Chroma (2025) document it independently across different models and tasks.
- The positional effect is real and asymmetric: the beginning and end of context are processed better than the middle, with documented performance drops exceeding 30%.
- Raw length is causal, not just content: whitespace tokens are sufficient to degrade performance (−48% for Llama at 30K tokens).
- Two mathematical models stand in opposition: exponential decay for Transformers in direct mode, polynomial decay for decomposed architectures (λ-RLM) — the difference is practically significant over very long contexts.
- The most robust operational solution is decomposition: reduce the size of subproblems submitted to the model, rather than extending the window.