rtk: A Single-Binary Rust Proxy That Trims 60-90% of the Command Output Your Agent Reads
rtk is an Apache-2.0 Rust CLI proxy from rtk-ai, shipped as one binary via Homebrew. It filters and compresses command output before it reaches an agent's context. The project claims 60-90% token savings. We review the design and the risks.
Over the past year a quiet shift has taken place in how coding agents spend their budget. The model now reads far more tokens than it writes. Each time an agent runs a shell command, whether it is a test suite, a git status, a directory listing or the build of a mid-sized project, the full output goes into the context window unchanged. A complete build log can run to thousands of lines, and only a handful of them may influence the next decision. The rest is billed per token, dilutes the model's attention, and pushes the session toward its context limit, which forces earlier summarization or truncation. rtk, an open-source project from the rtk-ai team, targets exactly this overlooked waste. Its tagline is blunt: a high-performance CLI proxy that cuts up to 90% of the bash output your agent reads. In terms of positioning, rtk is neither another agent framework nor another model gateway. It is a thin layer between the agent and the shell. Without it, the agent invokes a command and reads standard output directly. With it, the same command is routed through the proxy, which runs it, filters and compresses the result, and hands a shorter version to the model. The headline figure attached to the project is a token reduction of 60% to 90%, and the README phrases the upper bound as up to about 90%. These are the maintainers' own claims, and the ratio depends heavily on the command. Noisy, repetitive output such as dependency installation logs or verbose test progress offers a great deal of room for compression. Output that is already short offers almost none. Readers should treat the top of the range as a best case, not as an average.
The engineering choices behind the project are worth examining, because they are not accidental. An agent loop issues commands at a high rate, and a single task can trigger dozens or even hundreds of them. If the proxy depends on an interpreter or a heavy runtime, its startup latency is paid on every call and accumulates into a noticeable delay, and it adds environment dependencies that someone must maintain. A native executable with no external dependencies keeps the per-call overhead low, installs by being placed on the path, and behaves identically across containers, continuous integration machines and developer laptops. The repository ships a Homebrew formula, uses the permissive Apache-2.0 license, displays a security check badge from continuous integration, and offers its README in seven languages: English, French, Chinese, Japanese, Korean, Spanish and Portuguese. None of this is glamorous, but together it signals a project that expects real adoption rather than a weekend demo.
Every form of compression is a decision about what to discard, and that deserves direct attention. A human reader who misses one warning line in a long log usually loses little. An agent that misses a single critical error may continue from a false assumption, burn several more turns, and end up costing more than it would have without compression. For that reason, evaluating a tool like rtk by tokens saved alone is a mistake. The right metric is cost per successfully completed task. A sound procedure is to take a representative set of your team's real tasks, run each one several times with and without the proxy, and record completion rate, number of turns and total tokens. Failure paths deserve particular scrutiny: compiler errors, assertion failures, stack traces and exit codes must be preserved in full and not treated as noise. Teams should also keep an escape hatch that lets the agent read raw output when it appears confused. Observability is a related concern that is easy to overlook. Once a proxy sits in the path, the team needs to know what it removed, or post-incident analysis becomes guesswork. A good design leaves a trail for every compression: the original length, the compressed length, how many segments were folded, and, where possible, a way to recover the raw text for comparison. With that trail, token savings stop being a mysterious number and become an engineering metric that can be audited and tuned. In regulated environments there is a further question to settle before adoption, namely whether the full raw output is still retained locally for compliance and debugging. Zooming out, rtk reflects a broader movement in which context engineering is shifting from the prompt layer down to the tool layer. For the past two years, industry attention went to prompt compression, cache reuse and retrieval filtering. As agents became long-running loops, tool results turned into the dominant source of context, and trimming at the tool boundary is often cheaper and more predictable than summarizing afterward. The approach also complements prompt caching instead of competing with it: caching lowers the unit price of a repeated prefix, while a proxy reduces how much material enters the prefix in the first place, so the two effects can stack. The project is also a reminder that cost reduction does not always require a smaller model. Sometimes it only requires asking the model to read less of what it does not need.
Some restraint is still appropriate. As of this writing, the exact filtering rules, the range of commands rtk handles well, and the benchmark behind the 60% to 90% figure should be confirmed against the repository documentation and your own measurements. We have not independently reproduced those numbers, so they should not be written into a budget as fact. For teams that rely heavily on coding agents, the pragmatic path is to pilot the tool on a non-critical project, collect a week of token spend and task success data, and then decide whether to roll it out. If the data supports it, rtk is a low-cost way to obtain a visible reduction in spend. If it does not, the exercise will at least show you where your agents have been spending their money.
Sources
FAQ
What problem does rtk solve?
It reduces the tokens an agent spends reading command output. Build logs, test reports, directory listings and git output contain large amounts of text that does not change the agent's next decision, yet it is billed and fills the context window. rtk filters and compresses that output before the model sees it, and the project claims savings of up to roughly 90%.
Can compressing output make an agent miss a real error?
Yes, that is the central risk and the thing to test first. Check that failure messages, line numbers and exit status survive compression intact. Run your own task set with and without the proxy, compare success rate, turns and total tokens, and keep a way to bypass the proxy and read raw output.
Why a single Rust binary instead of a script or plugin?
A proxy runs on every command an agent issues, so startup cost is paid again and again. A native binary with no runtime dependencies starts quickly, is easy to ship through package managers such as Homebrew, and behaves the same on laptops, containers and CI machines. Actual speed figures should be measured on your own hardware.