What Ponytail does
Ponytail gives the agent a "lazy senior developer" rule set: efficient, not careless. Before writing anything it works down a short ladder (does this need to exist at all? Is it already in the codebase? Does the standard library or a native platform feature cover it? Can it be one line?) and stops at the first answer that works. The result is smaller, plainer code, produced with fewer output tokens.
It does not simplify the things that matter. Input validation, error handling that prevents data loss, security, and accessibility are explicitly exempt, and anything you ask for directly gets built in full.
Install in Claude Code
Ponytail is distributed as a Claude Code plugin from its own marketplace. Send these as two separate prompts; the install does not work if you paste both at once:
/plugin marketplace add DietrichGebert/ponytail
/plugin install ponytail@ponytail
The same two commands work in the Code tab of the Claude desktop app. The plugin runs two small Node.js lifecycle hooks, so node needs to be on your PATH; without it the commands still work, but Ponytail does not switch itself on at the start of each session. To remove it, run /plugin remove ponytail.
Install in Codex
codex plugin marketplace add DietrichGebert/ponytail
codex plugin add ponytail@ponytail
Then run codex, open /hooks, trust Ponytail's two lifecycle hooks, and start a new thread. The same install covers the ChatGPT desktop app once you restart it. In Codex the commands are skills, so you invoke them with @ (for example @ponytail-review).
It matters more on Codex than it used to. Since April 2026, ChatGPT Plus, Pro, and Business meter Codex by tokens rather than messages, so long answers drain the same 5-hour and weekly windows as input. See the Codex usage limits guide if you regularly hit the cap.
Install with the Headroom app
- Download the Headroom app and install it (drag to Applications on macOS, run the installer on Windows, install the .deb or AppImage on Linux).
- Open it once. It sets itself up and lives in the menu bar.
- Open the Add-ons screen and toggle Ponytail on. For Codex, enable the ChatGPT Codex connector as well.
- Start a new Claude Code or Codex session. Sessions that were already open do not pick up the change.
The app keeps Ponytail updated and one toggle covers both agents. The manual install is the better fit if you already manage your own plugins and want to pin the version you run.
Check that it is on
Start a fresh session and run /ponytail-help (@ponytail-help in Codex). If Ponytail is installed you get a reference card listing the levels and commands. If nothing happens, the session predates the install or the plugin did not load.
The behavioural check is more convincing: ask for something easy to over-build, like a cache in front of a function. With Ponytail on, the agent should reach for the smallest thing that works and say what it skipped, rather than hand you a configurable cache class you did not ask for.
Levels and commands
Switch levels with /ponytail lite, /ponytail full, /ponytail ultra, or /ponytail off; /ponytail on its own reports the current level. Full is the default and enforces the whole ladder. Lite is a lighter touch for exploratory work, and ultra cuts hardest. To change the default for every new session, set PONYTAIL_DEFAULT_MODE or a defaultMode field in ~/.config/ponytail/config.json.
/ponytail-reviewreviews the current diff for over-engineering and returns a list of things to delete./ponytail-auditdoes the same across the whole repository./ponytail-debtcollects theponytail:comments marking shortcuts it deliberately took, so they get tracked./ponytail-gainshows the project's benchmark results.
Where Ponytail fits
Over-built code costs twice: once in the output tokens to generate it, and again each time the agent re-reads and edits it. Ponytail cuts both. It does nothing about what the agent reads, which in most sessions is the larger number: tool output, logs, files, and earlier turns sent again every time. Headroom compresses that side locally before the request leaves your machine, typically 25-50% of a session's tokens. Caveman pairs with it too: Caveman shortens the agent's prose and leaves code alone, Ponytail shortens the code and leaves prose alone. Headroom vs Ponytail vs RTK shows how the pieces divide the work.