Results
| Reduction | Tokens | At list prices | |
|---|---|---|---|
| Input (measured) | -34% | 3.50B removed of a 10.30B baseline | $19,568 |
| Output (estimated) | -33% | 132M avoided of a 406M baseline | $7,700 |
| Total | 3.63B | $27,268 |
These are the same Input and Output figures the app's dashboard shows each user for their own traffic, computed over all 183 users combined.
How the savings are produced
On every turn, Claude Code appends new content to the conversation: tool results, file reads, search output, test and build logs, JSON from MCP servers. The proxy routes each block to a content-aware compressor and drops what the model is unlikely to need. Dropped content stays on your machine, and the model can fetch it back with a tool call, so compression is reversible. Context already in Anthropic's prompt cache is left as sent, so cache hits are unaffected.
The input figure is measured: tokens removed, divided by tokens removed plus the new input that reached the provider. Context re-sent from earlier turns is excluded from both sides.
For output, Extra Headroom instructs the model to answer more concisely. Avoided output is estimated against a learned baseline of reply lengths for the same kinds of requests, and a small control group of conversations runs without the instruction. Shorter replies also complete sooner, since generation time scales with output length.
Distribution across users
Per user, the median input reduction for the week was 29.9%, and the median user saved $44.30 at list prices, about $190 a month extrapolated. A quarter of users saw 40% or more; one in ten saw over 50%.
| 25th percentile | Median | 75th percentile | 90th percentile |
|---|---|---|---|
| Input -21.8% | Input -29.9% | Input -40.5% | Input -50.7% |
The reduction increases with volume. Split into thirds by new input per active day, the heaviest users saw the largest reduction:
| New input per active day | Users | Median input reduction |
|---|---|---|
| Under 2.3M tokens | 61 | -27.2% |
| 2.3M to 6.6M tokens | 61 | -29.6% |
| 6.7M tokens and up | 61 | -31.4% |
Rate limits
Most of these users are on Claude Pro or Max, where the dollar figures translate into usage rather than spend. On 58.5% of user-days, Claude Code received at least one HTTP 429 from Anthropic. Tokens removed locally never reach Anthropic, so they never count toward those limits.
Related: what's eating the Claude Code weekly limit and why sub-agents make long sessions expensive.
Methodology
The desktop app syncs daily per-user totals: token and request counts, never prompt content. The sample is every user on 0.9.21 or later (the first build that reports new-input tokens) whose requests that week were mostly Claude Code and who sent at least $5 of input at list prices: 183 users and 686 user-days; 71 of them also used another coding agent. The results table pools tokens across users; the distribution tables use each user's own weekly rate.