In short
Agentic systems deployed in production — Claude Code, GitHub Copilot, LangChain — compress between 75 and 95% of accumulated conversational context. Reference academic benchmarks show that quality drops beyond 50–60% compression. Seven independent sources draw a clear bimodal distribution: industry on one side, research on the other. The explanation comes down to one word: decomposition.
In short: imagine a chef preparing a full menu. If they had to keep every ingredient, every gesture, every pan used from starter to dessert in their head, they would collapse before the main course. But if they break their service into independent stations — chopping / cooking / plating — each station can forget the previous ones. Aggressive compaction works exactly that way: forget what the next step does not need.
The surprising figure
When a tool like Claude Code reaches the limit of its context window (the amount of text it can “see” at any given moment), it triggers a compaction mechanism. Compaction summarizes old exchanges to free up space, while preserving information deemed essential. The compaction rate measures the proportion of context eliminated: at 95%, only 5% of the original text remains.
Here is what seven independent measurements show:
| System | Publisher | Compaction rate | Type |
|---|---|---|---|
| Claude Code | Anthropic | 95% | Industrial (agentic) |
| GitHub Copilot | Microsoft | 95% | Industrial (agentic) |
| LangChain / LangGraph | LangChain Inc. | 85% | Industrial (agent framework) |
| Goose | Block | 80% | Industrial (agentic) |
| ForgeCode | ForgeCode | 75% | Industrial (agentic) |
| Liu et al. | EMNLP 2024 | ~50% | Academic (monolithic benchmark) |
The boundary is clear. Above 75%: production tools. Below 60%: research benchmarks. None of the measured systems falls in between.
Two worlds, two measures
Why such a gap? The answer starts with a difference in protocol.
Academic benchmarks test compression on monolithic tasks: the model receives a long document, a portion is compressed, then questions are asked about the whole. This is the protocol of Liu et al. (“Lost in the Middle”) and Veseli et al. (COLM 2025). In this scenario, any piece of information may be needed at any time. Compressing beyond 50–60% means losing details the task might require.
Industrial systems operate differently. A Claude Code agent does not process a single long problem in one block. It decomposes: read a file, modify a function, run a test, verify the result. Each sub-task is relatively autonomous. The full history of the conversation — the previous 40 exchanges — is not necessary to execute the current step.
This is precisely what Roy et al. formalize in their work on λ-RLM: for decomposable tasks, accuracy remains constant regardless of history length. The model only needs the local context of the current sub-task, not everything that preceded it.
The mechanism: decomposition makes compaction viable
The reasoning is circular at first glance, but it holds: agentic systems can aggressively compress because their architecture breaks work into atomic steps.
A concrete example. A developer asks Claude Code to “refactor the authentication module”. The agent will:
- Read the relevant files (context needed: the request + the files)
- Identify the modifications (context needed: the read files + project conventions)
- Apply each modification (context needed: the plan + the target file)
- Run the tests (context needed: the test command)
At step 4, the agent no longer needs the raw content of the files read at step 1. Compaction can eliminate those details without affecting the quality of the current step. The required residual context — “which file I modified, why, which test to run” — fits in a few hundred tokens.
This mechanism converges with an observation we made in our own experiments: the factual recovery rate (the ability to retrieve information from context) remains at 100% on contexts from 1K to 44K tokens, as long as the information is present. It is not the volume that matters — it is the presence or absence of the relevant information for the current task.
In short: compaction does not degrade quality as long as you only throw away ingredients you no longer need. It becomes destructive the moment the next step needs a detail you forgot. The 75-95% viability threshold is therefore not a property of the LLM — it is a property of how the task is decomposed.
The limits of this architecture
Aggressive compaction is not free. Three documented risks merit attention.
Loss of inter-step information. When a decision made at step 2 depends on a detail observed at step 1, and that detail has been compressed in the meantime, the agent may produce an inconsistency. This is the “context rot” phenomenon — the progressive degradation of contextual coherence — formalized mathematically by Roy et al. as exponential decay.
Dependency on decomposition quality. If the decomposition is poorly done — if a sub-task actually requires the context of several preceding steps — the high compaction rate becomes destructive. The quality of compaction depends directly on the quality of decomposition.
Survivorship bias. We measure the thresholds of systems that work. Systems that attempted 95% compaction without an adequate decomposition architecture have likely failed — and do not appear in our data.
Key takeaways
- Industrial agentic systems compress 75 to 95% of context; academic benchmarks plateau at 50–60%.
- This bimodal distribution is explained by decomposition into atomic tasks: each sub-task requires less residual context.
- Aggressive compaction is viable in production because the architecture breaks up the work — not despite it.
- Academic benchmarks test monolithic tasks, where all information may be needed at any moment — a different scenario from agentic use.
- The main risk remains inter-step information loss, especially when decomposition is imperfect.