In short
The context window is the maximum amount of text a language model can read and use at once. It includes what you send — your questions, your documents, the conversation history — and what it generates in response. When this limit is reached, information disappears. In five years, this limit has grown from 2,000 to several million tokens, but having a large window is not enough: models do not exploit every part of it with the same attention.
In short: think of the context window as a desk. Everything the model can consult must fit on this desk at the moment it writes its response. If you bring it a book, you take up space; if you ask for a long answer, it has to keep room to lay out its writing. When the desk overflows, the oldest sheets fall off — without warning.
Explanation
What it is
Picture the model as an expert who can only read a fixed number of pages at a time. You can ask anything — but only about what is in those pages. When you add new pages, the earliest ones drop off the stack. The model has not “forgotten” them in any neurological sense: they simply never existed in its field of view.
This limit is measured in tokens — units of text that roughly correspond to word fragments. In English, 1,000 tokens represent about 750 words. A standard novel runs about 100,000 words, i.e. 130,000 to 150,000 tokens.
The context window always covers both sides of the exchange: what you send (prompt, documents, history) and what the model generates (response). Both consume space.
The evolution: from 2K to 2M tokens
The progression is striking over five years.
| Model | Context size |
|---|---|
| GPT-3 (2020) | 2,048 tokens |
| GPT-3.5-turbo (2023) | up to 16,385 tokens |
| GPT-4 (2023) | up to 32,768 tokens |
| Claude 2 (2023) | 100,000 tokens |
| Gemini 1.5 Pro (2024) | 1,000,000 tokens |
| GPT-4.1 / GPT-5 (2025) | 1,000,000 tokens |
| Claude Opus 4.5 (2025) | 200,000 tokens |
In 2020, the 2,048-token limit made it impossible even to process a slightly long blog post. In 2025, a one-million-token window makes it possible to load a full encyclopedia, entire codebases, or months of exchanges in a single pass.
This evolution is not technically free: the compute complexity of a transformer grows as O(n²) with context length. Doubling the window quadruples compute. Recent advances (sparse attention, improved positional encoding) contain this cost, but the constraint remains real.
In short: going from 2K to 1M tokens in five years is a 500× factor. But the compute cost has been multiplied by 250,000 (quadratic growth). Progress comes as much from algorithmic optimizations as from raw GPU growth. Loading 1M tokens per request on a Claude Opus costs several dollars per call — hence the persistent interest in RAG architectures.
Concrete examples
Long documents
A 128,000-token window — common on current models — allows processing about 100,000 words in a single interaction. That is a medium novel, a complete due-diligence report, or several years of meeting notes. Before 2023, this kind of analysis required chunking the documents, processing them piecewise, and manually recombining the results.
Long conversations
In a multi-turn conversation, the model remembers nothing between sessions. It reconstructs its “memory” at every turn by re-reading the full history stored in context. When that history exceeds the window, the earliest exchanges disappear — and with them the initial instructions, the decisions made, and the constraints set.
RAG (Retrieval-Augmented Generation)
When a RAG system retrieves documents relevant to a question, it injects them directly into the model’s context window. A larger window allows injecting more documents at once, improving retrieval precision, and giving the model more context for a more nuanced answer. The window size is therefore a direct architectural constraint for any RAG system.
The “lost in the middle” problem
Having a large window does not guarantee that the model exploits everything in it. A 2024 study in the Transactions of the Association for Computational Linguistics (Liu et al., Stanford / University of Washington) documented a phenomenon called “lost in the middle”: models exploit information placed at the start and end of the context better than information placed in the middle.
On GPT-3.5-Turbo, performance could drop by more than 20% when the relevant information was in the middle of a long context — sometimes below the performance of a model that had been given no document at all. The shape of the curve is a U: both ends are exploited, the middle is partially ignored.
More recent work documents a related phenomenon called context rot: a performance degradation correlated with the raw growth of input tokens, independent of the quality of the retrieved documents.
In short: having a 200,000-token window does not mean the model uniformly processes 200,000 tokens. The more central the information is in the prompt, the more likely it is to be neglected. Practical rule: place critical instructions at the beginning and end of the context, and keep useful information at the top of the document.
Decision matrix — which context strategy to adopt
| Usage context | Recommendation | Why |
|---|---|---|
| Single long document (report, short book) | Native long context up to 50% of the window | Beyond, measured degradation ≥ 20% (Chroma 2025). |
| Large and evolving corpus | RAG with top-k relevant chunks | ~100× lower cost, higher precision when corpus > window. |
| Long conversation (persistent assistant) | Rolling summary + selective injection | Preserve initial constraints at the beginning without saturating. |
| Critical information to place | Beginning OR end of context | U-shaped positional bias, middle neglected. |
| Long generation expected (report, code) | Reserve 30-40% of the window for output | Generation also consumes tokens — risk of truncation otherwise. |
What to remember
- The context window sets the limit of what a model can process in one interaction. What falls out is lost — the model has no further access to it.
- It has grown from 2,000 tokens in 2020 to over a million in 2025, but this growth comes with a quadratic compute cost.
- A large window does not compensate for poor information placement: models process the start and end of the context better than the middle (“lost in the middle”).
- In RAG systems, window size is a direct architectural constraint that determines how many documents can be provided to the model in one pass.
- Filling the window is not always a good strategy: too much context can dilute relevant information and degrade performance (context rot).