oh-my-pi Deep Dive: Wiring the IDE into a Coding Agent, and Why the Edit Format Decides Model Success
oh-my-pi (omp) is a coding agent from Stencil Labs, forked from Pi, built around one idea: wire the IDE into the agent. It ships 31 built-in tools, 14 LSP operations, 28 DAP operations and a Rust core of about 80,000 lines. The author claims that tuning the edit format per model lifts Grok Code Fast 1 from a 6.7% to a 68.3% pass rate. This report covers the architecture, the persistent Python and Bun kernels with tool callbacks, and the limits. The benchmark figures are self-reported and need your own verification.
Positioning: an agent with the IDE wired in
oh-my-pi, invoked as omp on the command line, is maintained by can1357 at Stencil Labs. It is a fork of Pi, the open-source project by Mario Zechner. Its tagline is short: "A coding agent with the IDE wired in." The README quotes these scale figures: more than 60 model providers, 31 built-in tools, 14 LSP operations, 28 DAP operations, and about 80,000 lines of Rust in the core. These numbers come from the project itself and we did not verify each one. They still show the intent. omp does not want to be a model with a terminal. It wants the agent to have the full set of abilities a developer expects from an IDE.
The stack is TypeScript, Rust and the Bun runtime. It needs Bun 1.3.14 or newer and runs on macOS, Linux and Windows. The README also notes a policy trial: pull requests are open to everyone for now. Before, a contributor needed a vouch. The vouch system may return depending on how the trial goes.
Core architecture: three layers on top of Pi
The README suggests three layers. The first is the agent loop and terminal interface inherited from Pi. The second is what omp calls batteries: many built-in tools, model adaptation and prompt tuning. The third is the language-service layer that matches an IDE, namely LSP and DAP.
LSP, the Language Server Protocol, gives go-to-definition, find-references, diagnostics, rename and similar operations. The README says: "everything your IDE knows, the agent knows." With 14 LSP operations, the agent does not need to guess symbol relations with grep. It can ask the language server. DAP, the Debug Adapter Protocol, is rarer in agents. Its 28 operations let an agent set breakpoints, step through code and inspect variables. That replaces the crude loop of adding log lines and re-running with a real debug session. For an automated agent, these two protocols turn a text editor into a development environment.
The Rust core handles the performance-sensitive work. The README stresses search speed with lines such as "fastest in the west" for grep, and says searches return instantly. It does not publish implementation detail or search benchmarks, so we make no claims beyond that.
How it works: the edit format is the bottleneck
The central argument behind omp is that the harness, not only the model, limits quality. The author published a post titled "The Harness Problem" on 12 February 2026, and the README links to it. The idea: the same model performs very differently inside different harnesses, and the edit format matters most. If a model emits a diff in the wrong shape, the tool rejects it. The agent then enters a retry loop. Retries burn tokens and pollute the context. omp tunes tools and prompts for each model so that edits land on the first attempt. The README table lists four examples: - Grok Code Fast 1: pass rate rises from 6.7% to 68.3%, roughly tenfold. The author credits an edit format that stops "eating the model alive."
- Gemini 3 Flash: 5 percentage points above str_replace. The author says this beats Google's own best attempt at the format.
- Grok 4 Fast: 61% fewer output tokens, because the retry loop on bad diffs disappears.
- MiniMax: 2.1 times the pass rate, with the same weights and the same prompt.
Please note: these are the author's own figures. The test set, the sample size and the baselines are in the blog post, not in the README table. Treat them as leads worth checking, not as settled results. The underlying effect is plausible and worth knowing: the design of the edit tool can change how useful a given model is. The README lists a few more design points. The `read` tool returns summarized snippets instead of dumping whole files, with tuned defaults and a good selector hit rate. The `prompts` are adjusted relentlessly for each model.
Code execution with tool-calling
The first feature the README presents is code execution with tool-calling. Most harnesses give the agent a Python sandbox and stop there. omp runs a persistent Python kernel and a Bun worker. Either kernel can call back into the agent's own tools, such as read, search and task, over a loopback bridge.
In the README example, the agent loads a CSV with tool.read from inside Python, then charts it from JavaScript, and never leaves the cell. The value is clear. Intermediate data stays in the kernel. It does not have to flow back through the model context at each step. For large files and multi-step analysis, this saves context and matches how a person works in a notebook. The risk sits in the same place. Code execution plus tool callbacks widens the attack surface, so an enterprise should review the permission boundaries before wide use.
Installation and ecosystem
Install options are broad: a curl one-liner for macOS and Linux, Homebrew, a global Bun install (marked as recommended), Nix, Windows PowerShell, and pinned versions through mise. Nix users get a flake with packages, overlays, nixosModules and homeManagerModules. A Home Manager configuration can install omp and own its settings declaratively, for example `settings.startup.quiet = true`. On Alpine, install libstdc++ and libgcc first, because the prebuilt musl binary links them dynamically.
omp generates its own shell completions for bash, zsh and fish from live command and flag metadata, so they do not drift from the real CLI. Model names for `--model`, `--smol`, `--slow` and `--plan` complete against the bundled model catalog. `--resume` completes against sessions on disk. The flag names suggest that omp can use different models for different task stages, but check the official documentation for the exact routing rules.
Impact for developers and enterprises
For an individual developer, the appeal is a tool that works out of the box and stays open all the way down. With more than 60 providers, switching models is cheap. The per-model tuning matters most for teams that mix several vendors.
For an enterprise, the Nix and Home Manager support and the pinned-version install help with reproducible environments. The LSP and DAP integration lets the agent reuse the language-service setup a team already has.
Limits and outlook
First, most performance figures are self-reported and lack third-party replication. Second, 31 tools, two code kernels, LSP and DAP add a large surface area. Learning cost and security review cost are both real. Third, as a fork of Pi, omp must keep up with upstream, which is a long-term maintenance burden. Fourth, the pull request policy is still a trial, so community governance may change.
Likely directions include more per-model tuning, deeper debugging support and tighter permission control. Our advice is practical. Read "The Harness Problem" first. Then run omp on your own codebase with the models you already use, and test the percentages against your own tasks.