Zeroshot: The Agent That Writes the Code Never Approves It, Using Independent Review and Bounded Repair
Zeroshot turns a software goal into an explicit multi-agent graph. One agent implements, independent agents review, failures go into a bounded repair loop, and nothing ships until all the graph's checks pass. The implementing agent never approves its own work. It does not replace Claude Code, Codex or Copilot. It runs one of them as the worker and as the reviewers. Version 8 moves to a native binary shipped through npm, and you can bring your own topology and save it as a profile. The README publishes no benchmark data, so its value rests on design logic for now, not on measured results.
The problem it targets
The most common failure of AI coding agents is not that they cannot write code. It is that they finish and then declare the work done.
When one model both writes and approves the code, it grades its own exam. Zeroshot rests on a single claim: the agent that writes the code should not be the one that decides it works. The team says it built the tool because it was tired of being gaslit by agents that called broken code ready.
Core architecture: an explicit multi-agent graph
Zeroshot turns a software goal into an explicit multi-agent graph. One agent implements. Independent agents review. When a review fails, the work goes back into a bounded repair loop. Nothing is delivered until the graph's checks pass. The implementing agent never approves its own work. That is the hard rule of the design.
The positioning matters as much as the mechanism. Zeroshot does not replace Claude Code, Codex or GitHub Copilot. It runs one of them as the worker and as the reviewers. So it is not another coding model. It is an orchestration and accountability layer that wraps the agents you already use. You can use the built-in graph or bring your own topology: add reviewers, tests and repair loops, then save the setup as a profile for the next task.
How it works in practice
Installation is one command: npm install -g @the-open-engine-company/zeroshot. The installer needs Node.js 18 or newer. It installs a verified native binary for Linux x64 or arm64, macOS x64 or arm64, or Windows x64. It also installs one Zeroshot skill for Codex, GitHub Copilot and Claude Code at user scope. For local execution you install and sign in to Codex, Claude Code or GitHub Copilot. Zeroshot can reuse that harness's existing login, including subscription-backed sessions. A run needs two JSON files. input.json holds the task, for example: add JSON output to the status command and cover it with focused tests. runtime.json names the runtime: the harness (such as codex), the provider (such as openai), the model and the effort level (such as high). Then you call zeroshot run with --template software-change and --uniform-runtime-config runtime.json. From its name, the uniform option applies one runtime configuration across the agents in the graph. Add --validate-only to check the graph, the runtime configuration and the input before anything starts.
One fact deserves a warning. The worker edits the current Git worktree. The documentation therefore tells you to start in a clean worktree meant for that task.
Why independent review is a sound engineering choice
This section is our analysis, not data published by the project. First, separating the implementer from the verifier breaks the self-confirmation loop. The implementing agent carries a context full of its own reasons for why the change is right. A reviewer without that history is more likely to see what was missed. Second, the repair loop has a bound.
An unbounded cycle of review, fix and review again can burn quota without end. A bound turns cost and time into predictable quantities. Third, a topology that is configurable data rather than a fixed pipeline lets a team match the graph to the risk. A one-line documentation fix can use a light graph. A payment change can use several reviewers plus tests.
Performance and cost
We want to be plain here. The material we read contains no published benchmark numbers. We cannot quote a pass rate, a defect-catch rate or a latency change. The video in the README is labelled as scripted, so it is an illustration, not evidence.
The cost structure is clear, though. More agents mean more model calls and longer wall-clock time. If a local run reuses a subscription login, the marginal cost lands on that subscription quota. For Zeroshot Cloud pricing, see the product site at zeroshot.sh.
Impact on developers and enterprises
For an individual developer, Zeroshot turns the manual habit of asking another agent to look again into one repeatable command. For a team, it offers a verification layer that is not tied to one model vendor. The layer underneath can be Codex, Claude Code or Copilot.
The project uses the MIT license and ships through npm. The repository carries badges for build, coverage and documentation pipelines, which suggests it is maintained like production software. The project also claims stars from engineers at several large technology companies. That is the project's own statement, and we did not verify it independently.
Limitations and challenges
First, there is no public benchmark, so the value claim rests today on design logic. Second, independent review is only as good as its independence. If the reviewers and the implementer come from the same model family, they may share the same blind spots.
Third, a bounded repair loop means a task can still fail inside its bound, and users need a plan for that outcome. Fourth, because the worker edits the current worktree, running without a clean starting point is risky. Fifth, version 8 is a hard interface cutover. It replaces the Node.js runtime with a native binary, so users of older versions must migrate.
Where it may go
If the maintainers publish reproducible benchmarks, for example single-agent runs against review-graph runs on the same tasks with pass rate and total cost, the value of this kind of tool becomes measurable. Another direction to watch is profile sharing.
Teams that save tuned graphs for each risk level will slowly build an in-house standard for delivery. Either way, the principle that the author should not be the one to approve is becoming hard to avoid in AI-assisted programming.