哈佛研究发现AI编码智能体生成更多代码但未增加软件产出
AI coding agents generate more code, but not more software
哈佛大学Fiona Chen与James Stratton利用Jellyfish的汇总工程数据发现,企业使用AI编码工具后,几乎没有证据表明软件产出增加或用工减少。样本覆盖2021年至2026年3月、700余家软件公司、逾70万名员工和3亿条工作事件。编码阶段的效率提升被下游约束吸收,代码审查明显变长,pull request更常需要修订,审阅者留下的评论也更多。
Anyone who has even tangentially associated with computer programming knows that modern AI coding assistants and agents can be incredibly efficient at generating huge amounts of functional code. But coders making use of those tools also know better than to trust the accuracy of that code, meaning substantial effort needs to be spent reviewing any AI-generated output.
A recent study of actual coding practices across hundreds of firms finds that human code review forms a significant "bottleneck" for the overall efficiency of AI coding tools, resulting in "little evidence that firms increase software output or reduce employment" by using them. Any efficiency increased during the actual coding phase, the study authors find, is "absorbed by downstream constraints in the production process"; as "the code review process significantly increases in length, pull requests are more likely to require revisions, and reviewers leave more comments."
Cut once, measure twice
To come to these conclusions, Harvard University researchers Fiona Chen and James Stratton made use of aggregated analytics data from Jellyfish, which measures the granular output of engineering teams. That data encompasses 300 million individual "work events" (e.g., commits and pull requests) and issue management software data across more than 700,000 employees at over 700 relevant software development firms from 2021 through March of 2026.
来源:Ars Technica · AI · arstechnica.com