Cockroach Labs: Five Months Treating Bugs Like Patients and Coding Agents Like a Medical Team

Published · AI Daily — AI-assisted deep research, methodology & disclosure

Cockroach Labs ran a coding-agent pipeline modeled on a teaching hospital. In five months it handled over a million lines with seven reverts. Db2 support for MOLT took under two days and $4,172 in tokens: 164x faster, 38x cheaper than Oracle in 2024.

On a Wednesday evening in April, MOLT, the tool Cockroach Labs ships for migrating databases to CockroachDB, could not yet read from IBM Db2. By Friday afternoon, it could. Db2 is one of the first commercially available relational databases. It has a rich SQL dialect, a complex type system, and a native wire protocol. Supporting it meant a new schema converter, a new Fetch path for pulling rows out and loading them into CockroachDB, a new Verify path for comparing them once loaded, a vendored ANTLR grammar, a Docker image for CI, and more than ten thousand lines of test fixtures. When the team added Oracle support to MOLT in 2024, the equivalent work took nine months and cost about $160,000 in engineering time. The Db2 work took less than two days, and no human wrote any of the code. The token bill was $4,172. That is 164 times faster and 38 times cheaper.

The whole effort began with a single GitHub issue that described the request. A planning agent read it, decided it was too large to treat all at once, and decomposed it into fifteen sub-issues with an explicit dependency graph. The order was clear: first the foundation, then the type system, then the row iterator, then Fetch, Verify and Convert, and finally CI and test data. Two of the sub-issues were judged too large in their own workups, so the agents decomposed them again. Along the way, the agents filed about a dozen more issues against their own work: gaps in the fixtures, a type-mapping bug, and an isolation-level fix. No person found those problems. The pipeline found them, recorded them, and put them in the queue. Breaking the work apart and making its dependencies explicit is the first reason a task of this size could finish in two days.

The second reason is how the pipeline is organized. Cockroach Labs chose a teaching hospital as its guiding metaphor. Each bug is a patient, and each coding agent plays a role in a medical team. Triage Nurses receive and screen incoming work. They decide what the problem is, how large it is, and who should handle it. Fellows do the hands-on treatment, which means writing the fix. Review Attendings carry out rigorous review and guard quality. Discharge Nurses close the case and confirm that every check and every record is in place before the code merges. The metaphor is useful because it is not decoration. It takes a division of labor and a system of checks that software teams already know they need, and it fixes them in a workflow that everyone understands. A hospital does not let the person who treated a patient also sign off on the treatment. An agent pipeline will not respect that rule unless someone designs it in.

The third reason is a short list of hard disciplines. Agents must plan before they write code, and the plan itself goes through review, so that a wrong direction is stopped before tokens are spent on thousands of lines. Strict test coverage is a condition of merging, not something added afterward. Meticulous record-keeping means that every decision and every change leaves a trace, as a patient chart does, and that trace is the only way to understand later why something was reverted. Safety guardrails stand in front of the merge. These rules look plain. Yet they address exactly the places where agents fail most often: a misread goal, unearned confidence, and an output that nobody owns.

The most convincing measure of whether such a system is reliable is the number of reverts. Over five months, the pipeline processed more than a million lines of code and produced only seven reverts. That figure suggests that autonomous output and enterprise-grade reliability do not have to conflict, provided reliability is built into the structure of the system and is not left to the hope that a single model is clever enough. The same mechanisms support features such as the Migration Assistant. When intricate database migrations are built and maintained by agents, planning review, test coverage and complete records are the main defenses against costly regressions. A fair reading also needs caution. The numbers come from Cockroach Labs itself, and the comparison with the Oracle project, nine months and about $160,000, is the company's own estimate. They make a strong case study. They do not make a universal benchmark.

For the wider industry, the article changes the question. For years the common question was which model writes the best code. This case points to a more useful one: how do you organize a group of agents so that they divide the work, review each other, leave a record, and keep a clear chain of responsibility when something goes wrong? When the cost drops from nine months to two days, and from $160,000 to a little over $4,000, the bottleneck moves away from the act of coding. It moves to describing the requirement, decomposing the task, and accepting the result. Teams that want similar outcomes should first invest in decomposition skill, review process and test infrastructure. Chasing a stronger model comes after that, not before. The hospital metaphor is memorable, but the lesson underneath it is plain engineering: give every step an owner, write everything down, and do not merge until the checks pass.

Sources

FAQ

What roles does the pipeline use?

It borrows from a teaching hospital. Triage Nurses receive and screen problems, Fellows do the hands-on fix, Review Attendings run rigorous review, and Discharge Nurses finish the pre-merge checks and records.

How much faster and cheaper was Db2 than Oracle?

Cockroach Labs says Oracle support in 2024 took nine months and about $160,000 in engineering time. Db2 took under two days and a $4,172 token bill: 164x faster and 38x cheaper. These are the company's own figures.

How reliable was the system over five months?

It processed more than a million lines of code with only seven reverts. The key practices were planning review before coding, strict test coverage, meticulous records and safety guardrails before merge.