Markdown version of https://extraheadroom.com/what-is-headroom

# What is Headroom?

Headroom is a desktop app for macOS, Windows, and Linux that cuts the tokens Claude Code and ChatGPT Codex send to the model, so your plan's usage limits last longer and API bills shrink. It is built on the open-source Headroom CLI.

## Quick answer

Headroom routes Claude Code and Codex through a proxy on your machine. The proxy compresses tool output, logs, file reads, and older messages before the request goes to Anthropic or OpenAI: about 50% fewer tokens on the noisy items it compresses, typically 25-50% across a session. The compression is reversible, so the agent can retrieve the original whenever it needs the detail. The app also installs optional add-ons like RTK and Ponytail, records project learnings, and shows your usage windows and savings in the menu bar.

## How it works

Each request a coding agent makes carries far more than your prompt: file contents, command output, search results, and the whole conversation so far. Most of that is machine-generated and repetitive. Headroom compresses it locally, keeps the originals on your machine, and forwards the smaller request to your provider. When the model needs the full text of something it compressed, a retrieval tool fetches it back.

Savings depend on the work. Debugging from logs, shell-heavy tasks, and large JSON or HTML inputs compress well; short prompts barely change. Fewer tokens per task also means less inference energy; see [how Headroom makes AI coding more sustainable](/blog/sustainable-ai-coding).

## Subscriptions and the API

On a Claude Pro or Max plan, or a ChatGPT plan with Codex, you do not pay per token, so the savings show up as rate-limit headroom: the 5-hour and weekly windows last longer. It does not change your plan's price or the limits your provider sets. On the API, fewer input tokens is a smaller bill. The dashboard's dollar figures are API-equivalent costs; [how savings are measured](/docs/how-savings-are-measured) defines them and separates measured input compression from estimated output reduction.

## Add-ons and project learnings

Seven optional [add-ons](/docs/add-ons) each target another source of tokens: RTK (terminal output), MarkItDown (documents), Serena and Codebase Memory (code navigation), Context7 (library documentation), and [Ponytail](/ponytail-claude-code) and Caveman (shorter code and replies). [Project learnings](/docs/how-learning-works) write reusable commands and environment details into the agent's instruction files so it stops rediscovering them. The [tutorials](/docs/tutorials) show both in common tasks.

## Privacy

Compression runs on your machine. Your requests still go to Anthropic or OpenAI as before; Headroom does not need your prompts or source code on its servers. The [privacy policy](/privacy) covers accounts and telemetry.

## The app and the open-source CLI

The compression engine is the [Headroom CLI](https://github.com/headroomlabs-ai/headroom), an open-source (Apache 2.0) project created by Tejas Chopra. The desktop app is built on it with his endorsement and adds installation, automatic updates, agent connectors, add-ons, and a dashboard. The CLI is free; the app has a 7-day free trial with no credit card, then paid [plans](/#pricing). [Desktop app vs CLI](/headroom-desktop-vs-cli) compares them.

## Installing and removing it

The [install guide](/install-headroom-claude-code) covers both the app and the CLI, and the [quickstart](/docs/quickstart) shows how to connect an agent and confirm it routes through Headroom. Quitting Headroom removes its routing; open a new terminal afterward. For a full cleanup, use **Settings \> Uninstall Headroom** as described in the [uninstall guide](/docs/uninstall); deleting the app alone leaves integrations and local data behind. If an agent is doing the setup, point it at the [agent runbook](/docs/for-agents).
