In brief
Claude Code consumes between $2 and $210 per day depending on usage intensity, with cache reads representing more than 95% of tokens. The main optimization levers are prompt caching (×10 reduction on repeated tokens), model dispatch by task (Haiku for extraction, Sonnet for analysis), and active context management via /compact. This article details the settings, metrics, and anti-patterns identified after 30 days of intensive use.
Why monitor your consumption
Claude Code does not bill Max/Team users directly — but the tokens consumed determine rate limits, response speed, and the risk of throttling. On the API, the cost is direct: Opus 4.6 charges $15/MTok on input and $75/MTok on output. Thinking tokens (extended thinking) count as output tokens.
The problem is not the cost of an isolated prompt. It is the accumulation: a project with 10+ CLAUDE.md files, 8 MCP servers, and 2-hour sessions without /compact can consume 200M+ tokens per day. And a sub-agent launched without cache costs 10× more than an identical prompt with a cache hit.
Our observations over 30 days of intensive use: 4.1 billion tokens, of which 95.5% are cache reads. The median cost of a day with multi-agent orchestration is $158, versus $7 for a light direct-work day. The difference is a factor of 22×.
In short: 95% of consumed tokens are cache rereads at 10% of normal price. The whole optimization game consists in maximizing this ratio — therefore in avoiding triggers for cache rewrites (frequent system prompt changes, sessions too long, uncached sub-agents).
The metrics that matter
The 4 counters
| Counter | What it is | Why it matters |
|---|---|---|
| Input tokens | Text you send (prompt + context) | Negligible in volume (<0.1% of total) |
| Output tokens | Model response + thinking | The most expensive per token. Includes thinking. |
| Cache create | First write of a block to cache | 1.25× input price (one-time overhead) |
| Cache read | Re-reading an already-cached block | 10% of input price — the primary lever |
The ratio to watch: output / total. Based on our measurements, it oscillates between 0.1% and 0.5%. Below 0.1%, you are caching effectively but producing little. Above 1%, your output tokens (thus your thinking) dominate — reduce effort.
Cache is your best ally
Prompt caching automatically reuses the system prompt, CLAUDE.md files, and conversation history. A cache hit costs 10% of the normal price. Over our 30 days, cache reads represented 3.9 billion tokens out of 4.1 billion — without this mechanism, the cost would have been multiplied by a significant factor.
Activation conditions: the block must exceed 4,096 tokens for Opus/Haiku, or 1,024 tokens for Sonnet. Below that, failure is silent — no error, just full price. The cache expires after 5 minutes without a hit, but each hit renews the timer.
Documented pitfall: the --resume command caused a complete cache miss between Claude Code versions 2.1.69 and 2.1.90. Fixed since then. If you are using an earlier version, update.
In short: a cache hit is 90% savings on the cached portion. A silent cache miss (block too short, buggy version, parent → child) is 100% of the full price. The difference is worth periodically checking the cache hit rate in
/cost.
Optimization settings
ENABLE_TOOL_SEARCH — the most immediate gain
If you use MCP (Model Context Protocol) servers, this parameter is the first to check.
| Value | Behavior | Tokens at startup |
|---|---|---|
true (default) | Only tool names are loaded. Schemas on demand. | ~0 MCP tokens |
auto | Schemas loaded if < 10% of the window | Variable |
false | All schemas loaded upfront | Up to 70K+ tokens |
According to a community report, 8 MCP servers with false consume 70.5K tokens at startup — 35% of a 200K-token window, before a single question has been asked [NOT VERIFIED — single user measurement].
Configuration in ~/.claude/settings.json:
{
"env": {
"ENABLE_TOOL_SEARCH": "true"
}
}
Thinking and effort — the cost/quality slider
Thinking tokens are billed at the output rate ($75/MTok for Opus). This is the expense item that grows fastest under intensive use.
| Effort level | When to use it | Impact |
|---|---|---|
low | Mechanical extraction, lint, copies | Minimal thinking, fast response |
medium (default) | Writing, standard analysis | Good trade-off |
high | Multi-source synthesis, audit | Extended thinking, more expensive |
max (Opus) | Strategic judgment, meta-review | Maximum thinking budget |
The /effort low command at the start of a session reduces the thinking budget for the current session. For a permanent setting:
{
"effortLevel": "medium"
}
Practical advice: stay at medium by default. Switch to high only when reasoning is necessary (not for copy-paste or restructuring). Based on our observations, the majority of daily tasks (commits, targeted edits, code searches) do not benefit from extended thinking.
Model dispatch — choosing the right tool
Not all Claude models cost the same. Running mechanical extraction on Opus is wasteful.
| Model | Input/MTok | Output/MTok | Recommended use |
|---|---|---|---|
| Haiku 4.5 | $0.80 | $4.00 | Extraction, lint, scan, copy |
| Sonnet 4.6 | $3.00 | $15.00 | Writing, analysis, synthesis |
| Opus 4.6 | $15.00 | $75.00 | Complex orchestration, judgment |
For sub-agents, the model is configured in the agent’s frontmatter:
model: sonnet # or haiku, opus
Or globally via environment variable:
export CLAUDE_CODE_SUBAGENT_MODEL=sonnet
Based on our tests, Sonnet produces results equivalent to Opus on document extraction and standard analysis tasks. The gap only manifests on complex multi-step reasoning.
In short: three main levers to set before any regular use —
ENABLE_TOOL_SEARCH=true(frees 70,000 tokens of MCP context),effort mediumby default (high reserved for complex reasoning), model assigned to the task (Haiku for extraction, Sonnet for analysis, Opus for decision). These three settings cover ~80% of the optimization margin.
Costly anti-patterns
1. Long sessions without /compact
Context grows with each exchange. After 20+ turns, accumulated tokens weigh on every request. The /compact command summarizes history and frees up space.
Simple rule: /compact after each change of work axis. /clear when you change topics entirely.
Auto-compaction triggers at ~95% of the window, but by that point you have already been paying for an overloaded context for several turns.
2. Sub-agents without cache
This is the most costly trap in multi-agent orchestration. Each sub-agent opens its own context window. The parent’s prompt caching does not propagate to children — this is a documented behavior. Each sub-agent pays the full price for its initial context.
Concrete impact: based on our measurements, cost per sub-agent grew from $0.85 to $2.50 as prompts became more complex. Over a session with 40 sub-agents, the difference with/without hypothetical cache represents a significant multiple.
Mitigation: keep sub-agent prompts short and self-contained. Only inject the context strictly necessary for the task.
3. Oversized CLAUDE.md
A 500-line CLAUDE.md is reloaded at every conversation turn. If you have rarely-used instructions, move them to separate files loaded on demand (via skills or context files).
4. Opus for everything
The reflex to use the most powerful model for every task is natural but expensive. Opus costs 5× more than Sonnet on input and 5× more on output. For a reformatting, sorting, or data extraction task, Haiku at $0.80/MTok on input does the same work as Opus at $15/MTok.
5. No monitoring
Without a tracking tool, you do not know where your tokens are going. The most expensive days are not always the most productive.
Monitoring tools
/cost — quick diagnostics
Available since version 2.1.92, /cost displays the breakdown of the current session: tokens per model, cache hit rate, estimated cost. This is the first command to know.
/stats — usage patterns
For subscribers, /stats shows usage patterns over time.
ccusage — daily tracking
The ccusage tool (npx ccusage) provides a day-by-day table with input, output, cache create, cache read, total, and estimated cost. It is the reference tool for identifying anomalous days.
Analysis pattern: compare cost per day with the number of tasks completed. Based on our tracking, the most expensive days (>$150) systematically correspond to multi-agent orchestration sessions (40–200+ sub-agents). Light direct-work days cost $2 to $12.
/context — context X-ray
The /context command displays in real time what is occupying your context window: loaded files, history, system prompt. Use it when you suspect an overloaded context.
Quantified feedback — what a month of usage shows
Based on our tracking over 30 active days of intensive use (daily multi-agent orchestration):
| Metric | Value |
|---|---|
| Total tokens | 4.1 billion |
| Cache read | 95.5% of total |
| Median cost/day (orchestration) | $158 |
| Median cost/day (direct work) | $7 |
| Output/total ratio | 0.23% median |
| Most expensive day | $210 (91 sub-agents) |
| Cheapest day | $1.15 (Haiku only) |
The levers with the most impact:
- ENABLE_TOOL_SEARCH=true: immediately frees 70K+ tokens of MCP context
- Model dispatch: Haiku for mechanical extraction reduces cost per sub-agent by 5–10×
- Adapted effort:
lowormediumfor 80% of tasks,highonly when reasoning justifies it - Short sub-agent prompts: every token counts when there is no cache
- Regular /compact: keeps context below 50% of the window
Key takeaways
- Cache reads represent >95% of tokens: prompt caching is the most powerful optimization mechanism. Make sure it works (up-to-date version, blocks > 4,096 tokens for Opus).
- Cost comes from output, not input: thinking tokens (high/max effort) are billed at the output rate. Match effort to the task.
- A sub-agent does not benefit from the parent cache: keep prompts short and self-contained.
- Monitor to understand:
/cost,/context, andccusagemake visible what is usually opaque. A $150 day is not a problem if it produces 100 structured files. - Model dispatch is an underused lever: Haiku to extract, Sonnet to analyze, Opus to decide.