caveman: Cutting Coding-Agent Token Costs by Talking Like a Caveman
caveman is an open-source toolkit for coding agents with three parts: a skill that makes the model answer tersely, a local proxy that compresses tool output such as logs and diffs while keeping byte-exact originals retrievable, and middleware for developers who build their own agents. The headline claim of a 65% token cut comes from marketing. JetBrains measured 8.5% fewer output tokens across 86 tasks with the skill alone, with no detectable quality change. The proxy's 33.2% input saving comes from the maintainers' own benchmark. This analysis separates these numbers and explains the architecture.
Background and Problem Definition
Coding agents bill by the token. Every token costs money, and an agent pays twice: once for what it writes and once for what it reads. The writing is the reply. The reading is everything the tools return, such as logs, test output, JSON files, diffs, and search results. Both sides add up on the invoice. caveman, a project by JuliusBrussee on GitHub, attacks both sides. It began in April 2026 as a joke. The README reports a number one ranking on Hacker News, and the repository now has more than 109,000 stars. Those figures are the project's own, and popularity does not show that the technique works. The measurements do that, and they are the subject of this article.
The problem statement is simple. Most agents write like cover letters and read like firehoses. The first half is a tone problem: the model pads answers with preamble, restatement, and throat-clearing. The second half is a plumbing problem: tool output is raw, verbose, and mostly redundant for the question at hand. The README headline says the project "cuts 65% of tokens". That figure comes from marketing copy. The measurements that exist are narrower, and in one case they come from the project itself. This article separates them, because a reader who merges them will overestimate the saving.
Architectural Core and Technical Principles
caveman has three components. Each one installs on its own. The skill is a rule file. It instructs the agent to answer in terse, caveman-like prose. The README states that it works across more than 30 agents, including Claude Code, OpenAI Codex, Gemini CLI, Cursor, and Windsurf. The install command is `npx skills add JuliusBrussee/caveman -g`. The skill changes only the model's own prose. It does not touch tool output. The proxy runs on the user's machine, between the agent and the model provider. Before a request leaves, the proxy compresses tool results such as logs, JSON, diffs, and test output. The original is stored byte for byte in a local SQLite database, with a recovery handle. When the model needs the full text, it calls a `caveman_retrieve` tool to fetch it. The README says the proxy passes authentication through byte-exact, so the provider still receives the request credentials unchanged.
The middleware serves developers who build their own agents. It wraps a single call in LangChain, the Vercel AI SDK, the OpenAI client, or the Anthropic client. Large tool results are replaced by shorter copies before the model sees them. The conversation history keeps every original byte, so the model can read the full version when the short one is not enough. The design rests on one principle that all three share: compression must be reversible. The shortened text is a view, and the original stays reachable through a handle. This matters. A lossy summary chosen by a heuristic can silently remove the one line the model needed. Reversibility turns that silent error into a visible extra tool call. The split of responsibilities is also informative. The skill addresses output, the proxy addresses input, and the middleware carries the input idea into custom code. Each piece can be adopted alone, so a team can start with the cheapest one.
Practical Evaluation and Applications
Three sets of numbers are on the table. They measure different things, so they should not be added together. The first comes from JetBrains. The lab ran a paired A/B test on 86 real coding tasks, using Claude Code 2.1.200 and the skill alone. The proxy was not part of the test. Output tokens fell by 8.5 percent, and cost fell by roughly 10 percent. A sign test over the 18 non-tied tasks gave p = 0.82. The average task score moved from 0.326 at baseline to 0.311 with the skill, a gap of 0.015. On this evidence quality did not change in a measurable way, but the saving was small. The second comes from an Adobe Research paper, CAVEWOMAN, which the README cites. The paper reports that output-side caveman style cut realized cost to 1.4 to 2.4 times the baseline per model, and up to 3 times in the best case. The figures are the paper's. This article has not reproduced them. The third comes from the project itself. A 54-run Claude Code benchmark found 33.2 percent fewer input tokens with the proxy, across six workloads. All 18 answer checks passed. The case-clustered 95 percent interval runs from 14.6 to 48.5 percent. The README also says the raw harness artifacts are not in the checkout, so the result is a pinned report rather than a public reproduction. The authors ran the test themselves, and readers should weigh that.
Two more points deserve attention. First, telemetry. The command-line tool sends usage statistics, including a random install ID and the user's IP address, and this is on by default. The `caveman telemetry off` command turns it off. The skill alone sends nothing. Teams with data rules should check this before installation. Second, the comparison with Headroom and RTK. The README quotes JetBrains data that puts RTK at a 7.6 percent median cost increase per task at low reasoning effort. That is a claim about a competitor, drawn from someone else's test. Where does it fit? caveman suits high-volume agent sessions where the reply is routine and the tool output is large. Log triage, test debugging, and repository exploration are good examples. It fits poorly where wording carries legal, safety, or product weight. The README makes the same point: security warnings and confirmation prompts come back in full sentences. A shorter reply is not automatically a safer one. The sound method is to measure before adopting. The `caveman trial` command runs a session with and without caveman and reports the difference. A percentage measured on another codebase says little about yours.
Industry Impact and Outlook
caveman changed the conversation more than the technique. It made a billing fact visible. Teams tend to watch output tokens, because those are the texts they read. The project reminds them that reading is usually the larger share of the bill, and that the proxy targets exactly that share. The project also shows that a joke can carry a serious engineering idea. The tone stays playful. The mechanism underneath is a reversible compression layer between the model and its tools. That layer is not specific to caveman, and competing projects such as Headroom and RTK explore related approaches.
Three lessons follow for practitioners. First, measure the workload before trusting a percentage. Second, keep compression reversible, because a lossy cut that cannot be undone turns a model's decision into a permanent loss. Third, track quality next to cost, because a saving that degrades answers is not a saving. The open question is whether read-side compression holds up on longer sessions, more tools, and larger codebases. The published evidence covers short coding tasks and one project's own benchmark. It does not settle the question. An independent replication of the proxy result would be the most useful next step. The Apache-2.0 license, which covers the whole repository from version 3.0.0 on, makes that replication straightforward.
Sources
FAQ
What does the caveman skill actually change?
The skill is a rule file that makes the model write terse prose. It changes only the model's own replies and leaves tool output untouched, which is why its measured saving on output tokens is small.
Did JetBrains report a quality loss?
No measurable loss. The sign test over the 18 non-tied tasks gave p = 0.82, and the average task score went from 0.326 to 0.311, a gap of 0.015.