In short
A language model does not “read” a document the way a human would, from start to finish. It processes a sequence of tokens all at once, in a single operation whose memory cost grows quadratically with length. For years, this constraint limited models to a few thousand tokens. Today some reach one million. But the advertised length and the actually usable length are not the same thing — and the research documents this clearly.
In short: think of a long context as a cathedral-ceilinged library. The length advertised by a model (200K, 1M tokens) is the height of the shelves. The actually usable length is the height the librarian can reach without a ladder. The higher the ceiling rises, the wider the gap grows — and most recent “millions of tokens” are shelves you look at without grasping them.
The fundamental bottleneck
At the heart of the Transformer architecture lies the attention mechanism: every token must compare its content with that of every other token in the sequence. For a sequence of n tokens, that means n × n comparisons — a so-called quadratic complexity, written O(n²).
In practice: a sequence of 128,000 tokens requires roughly 16,000 times more GPU memory than a sequence of 1,000 tokens for this attention computation alone. For a long time, this constraint capped commercial models at around 4,000 to 8,000 tokens.
The first major workaround came not from a new architecture, but from an algorithmic optimization. FlashAttention (Dao et al., NeurIPS 2022) reformulates the attention computation to minimize data transfers between the GPU’s high-bandwidth memory (HBM) and its fast but limited on-chip memory (SRAM). Without approximation — the mathematical result is identical — FlashAttention enables processing sequences 2 to 4 times longer at constant memory. FlashAttention-2 (2023) achieves 50 to 73% of the theoretical peak throughput of an A100 GPU, compared with 25 to 40% for the original version.
In short: FlashAttention changed nothing about the mathematics of attention — it just better organized the back-and-forth between the different memory levels of the GPU. It is the equivalent of a cook better arranging their workstation: same ingredients, but far fewer wasted steps. Result: 2 to 4× longer at identical memory cost, with no precision loss.
Extending the window without retraining everything
Training a model natively on sequences of one million tokens is prohibitively expensive. Research therefore developed methods to extend the window of an existing model with minimal fine-tuning.
These methods all build on RoPE (Rotary Position Embedding), which has become the dominant positional encoding. RoPE encodes a token’s position via a rotation in the representation space. Its key advantage: it can be modified after training, unlike absolute learned position embeddings.
Position Interpolation (Chen et al., Meta, 2023): the problem with naive extension is that the model encounters positions it has never seen during training, which produces erratic attention computations. The fix: rather than extrapolating into unknown positions, compress the position indices to bring them back within the known range. With only 1,000 fine-tuning steps, LLaMA can thereby scale from 2,048 to 32,768 tokens.
YaRN (Peng et al., 2023) refines this approach by observing that not all RoPE dimensions play the same role. Some encode nearby positions (high frequencies), others distant positions (low frequencies). YaRN treats each dimension differently according to its frequency, and adds a correction factor on attention scores. Result: 10× less data and 2.5× fewer training steps than prior methods for comparable performance on extended contexts.
LongRoPE (Microsoft Research, 2024) pushes the extension to 2 million tokens via a progressive extension strategy: fine-tuning at 256,000 tokens, then a second pass to 2,048,000. A readjustment step preserves performance on short contexts, avoiding the regression typical of these methods.
Distributing the computation across multiple machines
For training on sequences of several million tokens, even FlashAttention is not enough: the full sequence does not fit on a single GPU. Ring Attention (Liu, Zaharia, Abbeel, UC Berkeley, 2023) solves this problem by distributing the sequence across multiple machines arranged in a ring. Each node processes a block of the sequence and exchanges key-value blocks with its neighbors, with communication overlapped behind the computation. This approach enables processing sequences proportionally longer to the number of machines used.
For fine-tuning, LongLoRA (Chen et al., ICLR 2024) proposes an economical alternative: replacing dense attention with shifted sparse attention during training, while retaining full attention at inference. With this method, LLaMA 2 7B scales from 4,000 to 100,000 tokens on a single 8-GPU node.
What benchmarks reveal
The simplest test for evaluating contextual memory is the “Needle in a Haystack” (Kamradt, 2023): hiding a specific piece of information at different positions in a long document, then asking the model to retrieve it. When this test was applied to Claude 2.1, which advertised a 200,000-token window, the recall rate reached only 27% on long contexts. Gemini 1.5 Pro (Google DeepMind, 2024) is the most documented exception: 99.7% recall at 1 million tokens on this test, and 99.2% at 10 million tokens in experimental mode.
But this test only measures the ability to retrieve a single isolated fact. The RULER benchmark (NVIDIA, 2024), which covers 13 more varied tasks, paints a less flattering picture: across 17 evaluated models, nearly all show significant performance drops as length increases, including those that perform near-perfectly on the previous test.
LLaMA 3.1 70B, trained with advanced extension techniques, shows an estimated effective context length of only 64,000 tokens in RULER tests — while the advertised window is far longer.
The main cause identified: during training, distant positions in the sequence are massively underrepresented in the data. The model simply did not learn to pay the same attention there. Announcements about context windows should therefore be read with care.
A related phenomenon, documented by Stanford (Liu et al., TACL 2024), is now called “Lost in the Middle”: on multi-document question-answering tasks, performance is highest when the relevant information appears at the beginning or end of the context. When it is in the middle, performance drops significantly. This U-shaped curve echoes the primacy and recency effects well known in memory psychology.
In short: the advertised window of a model is a technical ceiling, not a uniform quality zone. LLaMA 3.1 70B advertises 128K but collapses from 64K on RULER tasks. Claude 2.1 advertised 200K and dropped to 27% recall on simple extractions. Reading marketing numbers with a mental ×0.5 factor is a good approximation for the truly reliable zone.
RAG or long context? The question remains open
Given LLMs capable of reading very long documents, a practical question has emerged: do we still need RAG (retrieval-augmented generation)? RAG extracts relevant passages on the fly via a vector index and injects them into a short context, rather than loading the entire document.
A comparative study by Google (EMNLP 2024) on Gemini 1.5 and GPT-4 shows that long context outperforms RAG in raw performance on question-answering tasks. But RAG remains significantly cheaper, and for more than 60% of the tested queries, both approaches produce identical answers. The authors therefore propose a hybrid approach: use RAG by default and switch to long context only when the query demands it.
Other work nuances this picture. A 2025 paper (LaRA) concludes that no universally superior solution exists:
| Task | Advantage |
|---|---|
| Global overview, cross-document comparison, multi-document reasoning | Long context |
| Precise factual query, cost-constrained | RAG |
| Queries without access to relevant information | Equivalent (> 60% of cases) |
A more fundamental limitation was recently identified: lengthening the context degrades performance even when information retrieval is perfect (Yen et al., 2025). This result suggests that the sheer volume of tokens in context disrupts processing, independently of the ability to locate the information. The structural cause of this phenomenon has not yet been explained.
Decision matrix — when long context vs RAG
| Usage context | Recommendation | Why |
|---|---|---|
| Single coherent document for global analysis (book, report) | Native long context up to 50% of declared window | Global view > selective extraction; beyond, degradation > 20%. |
| Large heterogeneous corpus, targeted query | RAG with top-k chunks | ~100× lower cost, higher precision when corpus > window. |
| Multi-document comparison with cross-cutting reasoning | Long context (Gemini 1.5 Pro, Claude Opus 4) | Recall maintained on Gemini 99.7% at 1M tokens (Needle in Haystack). |
| Precise factual query, critical latency | RAG | Latency 100-500 ms vs several seconds for 1M long context. |
| > 60% of cases (standard queries) | Hybrid approach RAG by default + switch to long context if needed | Google EMNLP 2024 study: > 60% identical answers between RAG and long context. |
| Critical information to place | Beginning OR end of context | Lost in the Middle: loss > 30% in middle. |
What to remember
- The quadratic complexity of attention is the fundamental bottleneck of long context. FlashAttention circumvented it via an algorithmic optimization, without approximation.
- Extension techniques like YaRN and LongRoPE allow expanding the context window of an existing model with minimal retraining, by treating position dimensions differently according to their frequency.
- The advertised context length and the actually usable length differ systematically. RULER (NVIDIA, 2024) documents this gap across 17 models.
- The “Lost in the Middle” phenomenon (Stanford, 2023) shows that LLMs preferentially use information at the beginning and end of context — a structural property, not a bug.
- RAG and long context each have strengths depending on the task. Neither method dominates universally, and both can be combined adaptively.