GitHub Copilot and Developer Productivity: An Observational Dose-Response Analysis
AuthorsAlex Heilman, Alex Kyllo, Emerson Murphy-Hill
Resources
The paper finds that engineers write significantly more pull requests in weeks when they use GitHub Copilot heavily, suggesting the coding assistant can boost developer efficiency even after controlling for effort and confounders.
Key results
Microsoft Cloud+AI individual contributors in the panel
Study window for weekly engineer-level telemetry
Highest GitHub Copilot usage weeks vs zero-usage weeks under the main depth-based model
High-breadth Copilot weeks vs zero-breadth weeks in the robustness check
What the paper found
This Microsoft study, GitHub Copilot and Developer Productivity: An Observational Dose-Response Analysis, uses 43 weeks of telemetry from 16,223 engineers in Microsoft’s Cloud+AI organization to estimate whether GitHub Copilot improves output beyond simply being used by stronger or busier developers. The authors avoid an infeasible A/B test and instead fit a Poisson Pseudo-Maximum Likelihood model with engineer fixed effects, week fixed effects, and controls for coding time and browser time, so the estimand is an efficiency effect: more completed pull requests at the same measured effort. In the main depth-based specification, weeks with the highest Copilot usage show a 40.5% increase in completed PRs relative to zero-usage weeks, with a monotonic gradient across Low, Moderate, and High usage and evidence of diminishing returns at the top end. They then run seven falsification checks, including a non-coding Microsoft 365 Copilot placebo, a teammate PR placebo outcome, lag/lead timing tests, PR-size decomposition, PR-type decomposition, and an alternate breadth-based treatment definition; these largely rule out generic AI engagement, team-level shocks, within-week task reallocation, slicing into smaller PRs, and easier task mix as simple explanations. The breadth-based robustness check still shows a 54% increase at high usage, reinforcing the same qualitative conclusion: within-engineer Copilot-heavy weeks are associated with substantially higher PR throughput, though the paper emphasizes that this remains observational and likely understates any total effect if Copilot also saves time.
Original abstract
Does GitHub Copilot (GHCP) make engineers more productive, or do the engineers who use it more differ from those who use it less? And even within a single engineer, are GHCP-heavy weeks just busy weeks in which more of everything gets done? We study these questions using 43 weeks of data from 16,223 software engineers across Microsoft's Cloud+AI organization. Engineer fixed effects address the first concern by comparing each engineer against themselves rather than against other engineers, eliminating time-invariant differences in skill, role, and team. Active coding time and browser time then enter a Poisson Pseudo-Maximum Likelihood model with two-way fixed effects to address the harder, within-engineer confound: that GHCP-heavy weeks coincide with high-effort weeks. This defines our estimand as an efficiency effect: more pull requests completed at equivalent levels of coding time. Engineers are estimated to complete 40.5% more PRs in their highest GHCP usage weeks relative to their zero-usage weeks, holding measured development effort constant. The gradient is monotonic with diminishing returns at high intensity. Seven robustness and falsification tests target the remaining plausible alternative explanations (non-coding AI engagement, team-level shocks, within-week task reallocation, cross-week contamination, PR slicing into smaller units, shifts toward easier task types, and sensitivity to how the treatment is operationalized). Under an explicitly stated conditional-independence assumption, the within-engineer design estimates a tool-specific efficiency effect that is consistent with all seven robustness tests.
Read the original paperMore in Code Generation
Browse all 43 papers →Compact Documentation for Coding Agents: A Benchmark, an Optimizer, and Why It Does Not Transfer
Md Shohel Arman, Igor Molybog
Better code documentation can faithfully reconstruct software, but surprisingly does not necessarily help AI coding agents fix real issues when the source code is already available.
Groupwise Agentic Grading and Advantage Redistribution for Code Agent RL
Jinhao Dong, Liang Zhao, Zihao Yue, Wenhan Ma, Linghao Zhang, Lei Li, Shicheng Li, Yifan Song, Bowen Ye, Fuli Luo
GAGAR helps code agents learn not only to pass tests, but to produce cleaner and more targeted implementations by redistributing RL credit according to agentic quality judgments.
Reinforcement Learning from Intermediate Renders for Image-to-Code Generation
Omri Kaduri, Kate Feingold, Phillip Isola, Tali Dekel
IR4RL improves image-to-code generation by rewarding models for making useful visual progress at every intermediate rendering step.