Markdown version of https://extraheadroom.com/ponytail-claude-code

# Ponytail for Claude Code and Codex: how to install and use it

[Ponytail](https://github.com/DietrichGebert/ponytail) is an open-source plugin by Dietrich Gebert that makes Claude Code and Codex (now part of ChatGPT) write the least code that solves the problem. Most token-saving tools work on what the agent reads; Ponytail works on what it writes: smaller diffs, fewer output tokens, less to maintain. This page covers the two ways to install it, how to confirm it is on, and how to use its levels and commands.

## Quick answer

In Claude Code, send `/plugin marketplace add DietrichGebert/ponytail`, then `/plugin install ponytail@ponytail` as a separate prompt. In Codex, run `codex plugin marketplace add DietrichGebert/ponytail` and `codex plugin add ponytail@ponytail`, then trust its hooks under `/hooks`. Or install the Headroom desktop app and toggle Ponytail on in the Add-ons screen, which covers both agents and keeps it updated. Start a new session and run `/ponytail-help` to confirm. Levels are lite, full (the default), and ultra.

## What Ponytail does

Ponytail gives the agent a "lazy senior developer" rule set: efficient, not careless. Before writing anything it works down a short ladder (does this need to exist at all? Is it already in the codebase? Does the standard library or a native platform feature cover it? Can it be one line?) and stops at the first answer that works. The result is smaller, plainer code, produced with fewer output tokens.

It does not simplify the things that matter. Input validation, error handling that prevents data loss, security, and accessibility are explicitly exempt, and anything you ask for directly gets built in full.

## Install in Claude Code

Ponytail is distributed as a Claude Code plugin from its own marketplace. Send these as two separate prompts; the install does not work if you paste both at once:

```
/plugin marketplace add DietrichGebert/ponytail
```

```
/plugin install ponytail@ponytail
```

The same two commands work in the Code tab of the Claude desktop app. The plugin runs two small Node.js lifecycle hooks, so `node` needs to be on your PATH; without it the commands still work, but Ponytail does not switch itself on at the start of each session. To remove it, run `/plugin remove ponytail`.

## Install in Codex

```
codex plugin marketplace add DietrichGebert/ponytail
codex plugin add ponytail@ponytail
```

Then run `codex`, open `/hooks`, trust Ponytail's two lifecycle hooks, and start a new thread. The same install covers the ChatGPT desktop app once you restart it. In Codex the commands are skills, so you invoke them with `@` (for example `@ponytail-review`).

It matters more on Codex than it used to. [Since April 2026](/blog/codex-token-metering), ChatGPT Plus, Pro, and Business meter Codex by tokens rather than messages, so long answers drain the same 5-hour and weekly windows as input. See the [Codex usage limits guide](/codex-usage-limits) if you regularly hit the cap.

## Install with the Headroom app

1. Download the Headroom app and install it (drag to Applications on macOS, run the installer on Windows, install the .deb or AppImage on Linux).
2. Open it once. It sets itself up and lives in the menu bar.
3. Open the **Add-ons** screen and toggle **Ponytail** on. For Codex, enable the **ChatGPT Codex** connector as well.
4. Start a new Claude Code or Codex session. Sessions that were already open do not pick up the change.

The app keeps Ponytail updated and one toggle covers both agents. The manual install is the better fit if you already manage your own plugins and want to pin the version you run.

## Check that it is on

Start a fresh session and run `/ponytail-help` (`@ponytail-help` in Codex). If Ponytail is installed you get a reference card listing the levels and commands. If nothing happens, the session predates the install or the plugin did not load.

The behavioural check is more convincing: ask for something easy to over-build, like a cache in front of a function. With Ponytail on, the agent should reach for the smallest thing that works and say what it skipped, rather than hand you a configurable cache class you did not ask for.

## Levels and commands

Switch levels with `/ponytail lite`, `/ponytail full`, `/ponytail ultra`, or `/ponytail off`; `/ponytail` on its own reports the current level. Full is the default and enforces the whole ladder. Lite is a lighter touch for exploratory work, and ultra cuts hardest. To change the default for every new session, set `PONYTAIL_DEFAULT_MODE` or a `defaultMode` field in `~/.config/ponytail/config.json`.

- `/ponytail-review` reviews the current diff for over-engineering and returns a list of things to delete.
- `/ponytail-audit` does the same across the whole repository.
- `/ponytail-debt` collects the `ponytail:` comments marking shortcuts it deliberately took, so they get tracked.
- `/ponytail-gain` shows the project's benchmark results.

## Where Ponytail fits

Over-built code costs twice: once in the output tokens to generate it, and again each time the agent re-reads and edits it. Ponytail cuts both. It does nothing about what the agent reads, which in most sessions is the larger number: tool output, logs, files, and earlier turns sent again every time. [Headroom](/what-is-headroom) compresses that side locally before the request leaves your machine, typically 25-50% of a session's tokens. [Caveman](https://github.com/JuliusBrussee/caveman) pairs with it too: Caveman shortens the agent's prose and leaves code alone, Ponytail shortens the code and leaves prose alone. [Headroom vs Ponytail vs RTK](/headroom-vs-ponytail) shows how the pieces divide the work.
