The hype vs. the grind
If you've been following the AI hype cycle, you've probably heard that AI coding agents are about to make software engineers 10x more productive. Real talk: it’s giving major overpromise vibes. A massive new study from Harvard researchers Fiona Chen and James Stratton analyzed over 300 million work events across 700+ firms, and the results show that while we’re getting more code, we aren’t necessarily getting more software.
By the numbers
Once firms start using AI agents, the raw output is wild. We're talking a 30% jump in lines of code, 20% more commits, and 23% more pull requests. Sounds like a W, right? Wrong. When you look at actual finished features—the stuff that actually matters—the resolution rate hasn't changed at all.
The bottleneck is real
The reason for the slump is the "human-in-the-loop" problem. Because AI isn't perfect, human devs are spending way more time cleaning up the mess. The review process time has ballooned by 49%, and the number of comments left on pull requests is up 35%. Basically, developers are spending all their time playing cleanup instead of building new things. Lowkey, the efficiency gains you get from letting an AI write code are getting immediately eaten by the time it takes for a human to sanity-check it.
Why it matters
Right now, using AI agents is a total double-edged sword. Even though 80% of firms are trying to use AI to help with reviews, humans are still doing the heavy lifting on over 75% of the work. If firms want this tech to actually pay off, they need to figure out how to stop the review process from becoming a never-ending cycle of revisions. Say less—it turns out that shipping more code isn't the same as shipping better products.






