In short
When Claude Code launches a sub-agent via the Agent Tool, that sub-agent does not have 200,000 tokens of context. It has approximately 100,000. This is what our tests show reproducibly: auto-compaction triggers at around 94–96K tokens, placing the effective budget at roughly 99–101K. Another important point: each sub-agent has its own independent budget. Two parallel agents do not share a common pool — each starts with its own personal envelope.
In short: think of each sub-agent as a temporary employee handed a briefcase of 100,000 tokens to spend. The briefcase is personal, non-transferable, and locked at 95% full — at which point the sub-agent must summarize its memory to free space and continue.
Context
Claude Sonnet and Opus models advertise a 200,000-token context window in the official Anthropic documentation. With Max, Team, or Enterprise access, this window can reach up to 1 million tokens. These figures are correct for a standard interactive session.
But Claude Code operates differently. It consumes a portion of the context for its own system: the system prompt, tool definitions, MCP schemas, memory files. According to Claude Code documentation and several independent analyses, this structural overhead amounts to 30,000 to 40,000 tokens before writing a single line of prompt. And when a sub-agent is launched via the Agent Tool, a question arises: what budget does it actually have?
Before our measurements, the common assumption — present in some GitHub issues like #10212 — was that sub-agents shared the parent’s budget of around 200K tokens. This is not what we observe.
What we observe
Our tests cover two independent runs of a Sonnet agent launched via Agent Tool, across two versions of Claude Code (v2.1.75 and v2.1.79).
Measured budget: ~99–101K tokens per Sonnet sub-agent.
In the first run, the agent consumed 92,051 tokens without triggering auto-compaction. In the intensive test (repeated file reads over multiple passes), auto-compaction triggered after reaching approximately 93,991 tokens. Applying the observed ratio — compaction occurs at approximately 95% of the maximum budget — yields an estimated ceiling of around 99,000 tokens. Replication in v2.1.79 confirms: 95,936 tokens before compaction, giving an estimated budget of approximately 101,000 tokens. The delta between the two runs (+4.2%) is explained by a slight difference in source file sizes, not by a software change.
The 128K limit was not reached empirically.
Third-party sources sometimes cite a 128K-token ceiling for Claude sub-agents. Our measurements do not confirm this figure for Sonnet: compaction occurs well before that, around 95–96K, indicating an effective ceiling significantly below 128K. This 128K figure remains an unverified hypothesis — it may correspond to a different model or a theoretical API limit.
Budgets are independent — no shared pool.
This is probably the most useful finding in practice. In a two-agent parallel test, one agent consumed 70,092 tokens, the other 76,766 tokens — a combined 147K tokens — without either triggering auto-compaction. If the budget were shared (200K for the parent and all its children), two agents at 70K+ each could not coexist without a collision. That is not what we observe: each agent has its own envelope, independent of the parent and of other agents.
Auto-compaction works and is transparent.
When a sub-agent reaches the threshold (~95% of the budget), Claude Code automatically triggers a context compaction: old messages are summarized and the session continues without interruption. From our observations, this mechanism has been stable since version v2.1.79 — no crash, no silent truncation. The Anthropic documentation on compaction describes this behavior as intentional.
Beware of the total_tokens counter after compaction.
After a compaction, the total_tokens field returned by the API reflects the post-compaction state, not the historical cumulative total. In our tests, this counter displayed 23,470 tokens while the run had actually consumed more than 93,000 tokens before compaction. This figure cannot be used as a comparative metric between an agent that has compacted and one that has not.
In short: compaction is useful but it resets the odometer. If you compare consumption between two agents and one of them compacted along the way, you are comparing incompatible odometers. For token benchmarks, capture the value just before compaction or use external counters.
Implications
To design a sub-agent task, the useful budget is ~100K tokens, not 200K.
In practice, if you assign a Sonnet sub-agent a task that involves reading large files, multiple processing passes, or a detailed system prompt, you can plan for an envelope of approximately 100,000 tokens. Beyond that, compaction triggers. The session continues, but the historical context is summarized — which can affect coherence on tasks requiring a precise state across the full dataset.
Parallelism is genuinely parallel.
Since budgets are independent, launching 10 agents in parallel does not reduce each one’s budget. Each agent has its own ~100K tokens. The limiting factor is not the token pool — it is the complexity of what each agent is individually asked to do.
The v2.1.75 → v2.1.79 update did not modify the budget.
This version included a fix for premature compaction. From our measurements, the context budget for Sonnet sub-agents remained identical before and after: the fix improved stability, not the envelope.
Which sub-agent for which need?
| Task context | Recommendation | Why |
|---|---|---|
| Simple read-extract, < 30K input tokens | Haiku or Sonnet, no budget concern | Comfortable margin to compaction (~94K), no truncation risk |
| Multi-file analysis, 30-60K input tokens | Sonnet, watch the budget | Entering the buffer zone — compaction stays distant but historical context consumes |
| Task exceeding 80K predictable input tokens | Split into 2 parallel agents | Compaction otherwise certain → better 2 clean envelopes than one destructive compaction |
| Massive parallelism (≥ 10 agents) | Confirm Tier 2+ API | Budget is independent per agent, but the ITPM rate limit is shared |
Key takeaways
- The effective budget of a Sonnet sub-agent is approximately 100,000 tokens, not 200,000.
- Auto-compaction triggers at approximately 94–96K tokens (i.e., ~95% of the ceiling).
- Budgets are independent per agent: no shared pool between parent and children.
- The 128K limit has not been empirically confirmed for Sonnet — treat this figure as a hypothesis.
- The
total_tokenscounter post-compaction is not comparable to a pre-compaction counter: do not use it for cross-agent metrics. - These observations are stable across two runs and two versions of Claude Code (v2.1.75 and v2.1.79).