AI Writes Faster Than Humans Can Review: A Longitudinal Study of an Enterprise 2x Mandate
AuthorsHao He, Shyam Agarwal, Yegor Denisov-Blanch, Pavel Azaletskiy, Sanmi Koyejo, Bogdan Vasilescu
Resources
A large enterprise study finds that an AI coding mandate eventually doubled developer output, while shifting much of code review from humans to automation.
Key results
Developers analyzed in the longitudinal panel.
Non-bot pull requests analyzed across the observation window.
Fold increase in pull requests per active developer by April 2026 versus the pre-mandate baseline.
Fold gain predicted from adoption and accumulated use with full calendar-month controls.
Share of pull requests receiving automated review near the end of the study.
What the paper found
Hao He and colleagues at Carnegie Mellon University and Stanford University studied an anonymized, AI-forward enterprise after its CTO introduced a 2× productivity mandate, analyzing 802 developers and 196,212 pull requests with a staggered difference-in-differences design, developer fixed effects, and telemetry from Cursor and Claude Code. Per-developer throughput rose from 21.2 to 44.3 pull requests per month, a 2.09-fold increase by April 2026, but the gain accumulated gradually rather than appearing immediately: adoption produced an initial jump, while cumulative AI use accounted for most of a conservative 1.46-fold within-developer gain after full calendar-month controls. The study spans Sonnet 4.5, Opus 4.5, and Opus 4.6, yet model-generation effects could not be identified because releases coincided with a firm-wide usage ramp and an August 2025 placebo moved as much as real releases. Gains were broadly similar across individual-contributor through principal levels, concentrated in newer repositories, and not significant in legacy code. Review capacity became the bottleneck: per-reviewer load approximately doubled, human-review coverage fell from 89% to 68%, and automated review reached 84%. Merge and revert rates remained broadly stable, but AI-authored pull requests experienced roughly 20% longer human review latency. The authors conclude that an enterprise mandate acts mainly as a catalyst for adoption and learning, while shifting work downstream from code production to review and automation; because adoption was nonrandom, the exact causal magnitude remains bounded.
Original abstract
Enterprises increasingly mandate AI coding tools and report large productivity gains, yet longitudinal evidence on how such a mandate unfolds is scarce. In this paper, we present a quantitative case study of a documented enterprise "2x" mandate at a mid-sized, AI-forward company that has been committed to doubling merged pull requests per engineer since mid-2025. In a panel of 802 developers and 196,212 pull requests (January 2024-April 2026), per-capita throughput eventually doubled, reaching 2.09x the pre-mandate baseline in April 2026, among the largest gains reported from a field deployment of AI coding tools to our knowledge. A staggered difference-in-differences design links the within-developer share of this gain to AI adoption and to a further gain that grows with accumulated use, with the mandate acting as a catalyst rather than a direct driver. Because adoption and usage intensity were not randomly assigned, we read this evidence as strongly implicating an adoption-and-use channel rather than as exact causal attribution. The gain is broadly shared across seniority yet concentrated in newer code and not separable across model generations. Adoption also restructured code review around automation: per-reviewer load roughly doubled and automated review overtook human review, while merge and revert rates held steady.
Read the original paperMore in Code Generation
Browse all 43 papers →Compact Documentation for Coding Agents: A Benchmark, an Optimizer, and Why It Does Not Transfer
Md Shohel Arman, Igor Molybog
Better code documentation can faithfully reconstruct software, but surprisingly does not necessarily help AI coding agents fix real issues when the source code is already available.
Groupwise Agentic Grading and Advantage Redistribution for Code Agent RL
Jinhao Dong, Liang Zhao, Zihao Yue, Wenhan Ma, Linghao Zhang, Lei Li, Shicheng Li, Yifan Song, Bowen Ye, Fuli Luo
GAGAR helps code agents learn not only to pass tests, but to produce cleaner and more targeted implementations by redistributing RL credit according to agentic quality judgments.
Reinforcement Learning from Intermediate Renders for Image-to-Code Generation
Omri Kaduri, Kate Feingold, Phillip Isola, Tali Dekel
IR4RL improves image-to-code generation by rewarding models for making useful visual progress at every intermediate rendering step.