NTH

AI Writes Faster Than Humans Can Review: A Longitudinal Study of an Enterprise 2x Mandate

AuthorsHao He, Shyam Agarwal, Yegor Denisov-Blanch, Pavel Azaletskiy, Sanmi Koyejo, Bogdan Vasilescu

July 22, 2026 2 min read
Watch on YouTube
The one-line take

A large enterprise study finds that an AI coding mandate eventually doubled developer output, while shifting much of code review from humans to automation.

Key results

802
Developer panel

Developers analyzed in the longitudinal panel.

196,212
Pull-request dataset

Non-bot pull requests analyzed across the observation window.

2.09
Per-developer throughput

Fold increase in pull requests per active developer by April 2026 versus the pre-mandate baseline.

1.46
Conservative within-developer gain

Fold gain predicted from adoption and accumulated use with full calendar-month controls.

84%
Automated review coverage

Share of pull requests receiving automated review near the end of the study.

What the paper found

Hao He and colleagues at Carnegie Mellon University and Stanford University studied an anonymized, AI-forward enterprise after its CTO introduced a 2× productivity mandate, analyzing 802 developers and 196,212 pull requests with a staggered difference-in-differences design, developer fixed effects, and telemetry from Cursor and Claude Code. Per-developer throughput rose from 21.2 to 44.3 pull requests per month, a 2.09-fold increase by April 2026, but the gain accumulated gradually rather than appearing immediately: adoption produced an initial jump, while cumulative AI use accounted for most of a conservative 1.46-fold within-developer gain after full calendar-month controls. The study spans Sonnet 4.5, Opus 4.5, and Opus 4.6, yet model-generation effects could not be identified because releases coincided with a firm-wide usage ramp and an August 2025 placebo moved as much as real releases. Gains were broadly similar across individual-contributor through principal levels, concentrated in newer repositories, and not significant in legacy code. Review capacity became the bottleneck: per-reviewer load approximately doubled, human-review coverage fell from 89% to 68%, and automated review reached 84%. Merge and revert rates remained broadly stable, but AI-authored pull requests experienced roughly 20% longer human review latency. The authors conclude that an enterprise mandate acts mainly as a catalyst for adoption and learning, while shifting work downstream from code production to review and automation; because adoption was nonrandom, the exact causal magnitude remains bounded.

Original abstract

Enterprises increasingly mandate AI coding tools and report large productivity gains, yet longitudinal evidence on how such a mandate unfolds is scarce. In this paper, we present a quantitative case study of a documented enterprise "2x" mandate at a mid-sized, AI-forward company that has been committed to doubling merged pull requests per engineer since mid-2025. In a panel of 802 developers and 196,212 pull requests (January 2024-April 2026), per-capita throughput eventually doubled, reaching 2.09x the pre-mandate baseline in April 2026, among the largest gains reported from a field deployment of AI coding tools to our knowledge. A staggered difference-in-differences design links the within-developer share of this gain to AI adoption and to a further gain that grows with accumulated use, with the mandate acting as a catalyst rather than a direct driver. Because adoption and usage intensity were not randomly assigned, we read this evidence as strongly implicating an adoption-and-use channel rather than as exact causal attribution. The gain is broadly shared across seniority yet concentrated in newer code and not separable across model generations. Adoption also restructured code review around automation: per-reviewer load roughly doubled and automated review overtook human review, while merge and revert rates held steady.

Read the original paper

More in Code Generation

Browse all 43 papers →
02Code Generation

Groupwise Agentic Grading and Advantage Redistribution for Code Agent RL

Jinhao Dong, Liang Zhao, Zihao Yue, Wenhan Ma, Linghao Zhang, Lei Li, Shicheng Li, Yifan Song, Bowen Ye, Fuli Luo

GAGAR helps code agents learn not only to pass tests, but to produce cleaner and more targeted implementations by redistributing RL credit according to agentic quality judgments.

Read analysis