Headroom: Pre-LLM Context and Token Compressor for Coding Agents and RAG
Headroom is an open-source layer that compresses agent and RAG input before the LLM sees it. Its demo cuts 55,957 tokens to 24,340, about 56.5% fewer, and a FATAL log line survives byte for byte. Apache 2.0, on PyPI and npm. One demo is no benchmark.
Headroom, an open-source project from headroomlabs-ai, is a compression layer that sits in front of the language model. It targets a narrow and expensive problem. Coding agents and retrieval-augmented generation (RAG) systems push large amounts of raw material into the prompt on every turn: tool outputs, run logs, retrieved chunks, whole files. Most of that material is verbose and repetitive, and the facts that decide the next action often fit in a few lines. The hero graphic on the project page makes the point with one example. An agent prompt of 55,957 tokens is reduced to the 24,340 tokens actually sent to the model, a cut of about 56.5 percent, while a single FATAL log line at item 67 survives byte for byte. The example frames the real difficulty of the category. Saving tokens is easy. Proving that the discarded part holds no fatal detail is hard. To see why this matters, look at the cost structure of a coding agent. During a task the agent reads files, runs tests and scans build logs. The return value of every tool call is appended to the conversation history and is sent again on every later turn. Context therefore grows by accumulation: a three-thousand-line log read early in the session is billed again and again across dozens of turns, and it keeps occupying the window. Money is only one side of the cost. A longer context raises inference latency, and it dilutes the model's attention to material in the middle of the prompt. The industry has long observed that long contexts lose information placed mid-way. Stopping noise before it reaches the model can lower cost, latency and error rate together. That is the appeal of the pre-processing route. It needs no change of model and no rewrite of the agent. It adds one layer between the two.
The public project material gives two firm clues about the technical route. First, the project publishes a model on Hugging Face named kompress-v2-base. This suggests that compression does not rest on regular expressions and truncation alone, and that it includes a learned compressor. Second, the demonstration stresses that the critical line survives byte for byte. This suggests that lossless retention of key signals is an explicit goal, not a side effect noticed afterwards. On top of those facts, the following is our analytical inference, not a documented claim. A system of this kind usually has to route content by type. Logs, JSON, source code and prose each carry a different kind of redundancy. Repeated stack frames, hundreds of records with the same shape, stray whitespace and boilerplate can be merged heavily. Anchors such as the error level, the exception name, the file path and the line number must pass through unchanged. The README excerpt available to us does not describe these mechanisms in detail. The exact algorithms, thresholds and evaluation method should be taken from the official documentation and the code.
On the engineering side, the project takes a practical stance. It ships to both PyPI and npm under the name headroom-ai, which covers the two largest ecosystems for agent development, Python and JavaScript. The licence is Apache 2.0, which suits enterprise adoption. The documentation site offers a quickstart, the front page promises an install in about 60 seconds, and it lists compatibility notes for a range of agents. The project also prepares an llms.txt file and a full documentation bundle for AI readers. That detail belongs to its time: documentation is now read by agents too, so the team applies "optimise for machine readers" to its own pages first. The page also links to Headroom for Teams, which hints at team-oriented features beyond the open-source core. The source material does not say what form they take. Trendshift ranked the repository first of the day, which shows community attention. Attention is not quality, and the next section explains why that matters. Every lossy context compressor carries one basic risk: the compressor does not know what the downstream task will need. A warning that looks irrelevant today may be the clue that solves a problem three turns later. The FATAL line preserved in the demo is an obvious anchor. In real work, the critical fact is often quiet: a configuration value that changed without comment, or a small difference in the order of events in a log. For that reason, one official example cannot replace a benchmark. A team that adopts the tool should run its own comparison. Take the same set of real agent tasks, run it with compression on and off, and compare task success rate, mean token use, number of turns and latency, not only the compression ratio. Two engineering points deserve attention as well. First, compression changes the byte sequence sent to the model. It may interact with the provider's prompt caching, and the savings from compression may be partly offset by a lower cache hit rate. Second, the compression layer becomes a new point of failure and a new object of audit. When something goes wrong, the team must be able to trace what the model actually saw.
Taken as a whole, Headroom stands for a category that is taking shape inside agent infrastructure. Context engineering is no longer only the craft of writing prompts. It is becoming a reusable, measurable layer of middleware. Even if context windows keep growing, the limits of cost and attention remain, and a larger window only invites more rubbish to be poured in. A team that runs many coding agents or retrieval pipelines each day should test this direction with a small pilot. Enable it first on log-heavy tasks, keep a side archive of the full original text, measure with your own regression set, and only then decide whether to widen the scope. For the general reader, two things are enough to remember: the pair of numbers and the one FATAL line. The value of a compressor depends on whether it can guarantee that the most important line comes through untouched.
Sources
FAQ
What is Headroom and what problem does it solve?
Headroom is an open-source compression layer between an agent and the LLM. It compresses tool output, logs, RAG chunks and files to cut token cost and latency, and to keep irrelevant text from diluting the model's attention.
How strong is the compression in the headline example?
The demo cuts a 55,957-token prompt to the 24,340 tokens actually sent, about 56.5% fewer, and the FATAL log line at item 67 survives byte for byte. It is one official demo, not a general benchmark.
What should a team verify before adopting it?
Run your own real tasks with compression on and off. Compare success rate, mean tokens, turns and latency, check the effect on prompt caching, and keep an archive of the original text so you can trace what the model saw.