AI coding agents generate more code, but not more software
Harvard researchers Fiona Chen and James Stratton analyzed Jellyfish data: 300 million work events, 700+ firms, 700,000+ staff. Coding gains are absorbed by slower code review, and firms show little rise in software output or cut in employment.
For several years, the industry has worked from a simple assumption. If AI coding assistants and agents write code faster, software output should rise in proportion, and companies might even need fewer engineers. A new study, reported by Ars Technica, challenges that assumption directly. Harvard University researchers Fiona Chen and James Stratton examined real engineering data from hundreds of firms and found little evidence that companies adopting these tools increase software output or reduce employment. Code is cheaper to produce. Software, as a finished and shipped good, is not.
The empirical base of the study is substantial. The authors used aggregated analytics from Jellyfish, a company that measures the granular output of engineering teams. The data covers roughly 300 million individual work events, such as commits and pull requests, together with issue-management records. It spans more than 700,000 employees at over 700 software development firms, from 2021 through March 2026. To decide when each company adopted AI, the researchers combined directly measured AI usage with an analysis of GitHub activity. They also separated two product types. AI coding assistants help auto-complete code that humans mainly author. AI coding agents write and submit code largely on their own, based on prompts. The team then ran a difference-in-differences regression on key variables before and after adoption, staggered across firms. This design matters. It separates broad time trends and fixed differences between companies from the effect of the tools themselves, and so it comes closer to a causal reading than a simple comparison of users and non-users.
The most useful part of the paper is its explanation. According to the authors, any efficiency gained in the coding phase is absorbed by downstream constraints in the production process, and the largest of these is human code review. After adoption, the review process grows significantly longer, pull requests are more likely to need revisions, and reviewers leave more comments. Every working programmer will recognize the pattern. AI output often looks plausible and often runs, yet nobody can safely trust it without checking. The cost of establishing correctness therefore moves from the author to the reviewer. When one stage of a pipeline accelerates and its neighbor does not, total throughput is set by the slowest stage. This is the old logic of the theory of constraints, with one change: the bottleneck is now reading, not writing.
For the industry, the finding carries several lessons. First, measures such as lines generated, commit counts, or tool adoption rates can mislead. They sit on exactly the side of the process that the tools amplify, so they will look impressive even when delivery does not change. Second, companies that want a real productivity return cannot stop at buying a stronger generator. They must also invest in the review stage: better automated testing, static analysis, enforceable conventions, and agents that help verify work rather than only produce it. Third, the story that AI will let firms shrink engineering teams lacks support in this data. Within the study window, employment shows no clear decline. It is worth stating the limits of that claim. The study does not say that AI delivers no value. It says that firm-level indicators of output and headcount show little change, which is a different statement from saying that individual tasks are not faster.
The work also has limits that readers should keep in mind. The data ends in March 2026, while model capability, agent design, and team workflows are changing quickly. The review bottleneck may not be permanent. If future agents produce smaller, easier-to-verify changes, or attach tests and evidence that reviewers can check quickly, the cost of review could fall. Software output is also hard to measure, and commits and pull requests are only proxies for value delivered to users. Even so, the message is clear. In the age of agents, the scarce skill is no longer producing code. It is confirming, at acceptable cost, that the code is correct. The teams and tools that make verification faster and more reliable are the ones most likely to turn more code into more software.
Sources
FAQ
What is the study's main finding?
Harvard researchers Fiona Chen and James Stratton found little evidence that firms using AI coding assistants or agents raised software output or cut employment. Gains in the coding phase were absorbed by downstream constraints, mainly human code review.
Why does code review become the bottleneck?
After AI adoption, review takes noticeably longer, pull requests more often need revisions, and reviewers leave more comments. AI output cannot be trusted unchecked, so the cost of verifying correctness shifts from the author to the reviewer.
How reliable are the data and method?
The data come from Jellyfish: about 300 million work events across 700+ firms and 700,000+ employees, from 2021 to March 2026. The method is a difference-in-differences regression. Limits: the data stop in March 2026, and commits and pull requests are only proxies for output.