NTH

GitHub Copilot and Developer Productivity: An Observational Dose-Response Analysis

AuthorsAlex Heilman, Alex Kyllo, Emerson Murphy-Hill

July 5, 2026 3 min read
Watch on YouTube
The one-line take

The paper finds that engineers write significantly more pull requests in weeks when they use GitHub Copilot heavily, suggesting the coding assistant can boost developer efficiency even after controlling for effort and confounders.

Key results

16223
Engineers

Microsoft Cloud+AI individual contributors in the panel

43
Weeks

Study window for weekly engineer-level telemetry

40.5%
High usage PR increase

Highest GitHub Copilot usage weeks vs zero-usage weeks under the main depth-based model

54%
Breadth-based high usage PR increase

High-breadth Copilot weeks vs zero-breadth weeks in the robustness check

What the paper found

This Microsoft study, GitHub Copilot and Developer Productivity: An Observational Dose-Response Analysis, uses 43 weeks of telemetry from 16,223 engineers in Microsoft’s Cloud+AI organization to estimate whether GitHub Copilot improves output beyond simply being used by stronger or busier developers. The authors avoid an infeasible A/B test and instead fit a Poisson Pseudo-Maximum Likelihood model with engineer fixed effects, week fixed effects, and controls for coding time and browser time, so the estimand is an efficiency effect: more completed pull requests at the same measured effort. In the main depth-based specification, weeks with the highest Copilot usage show a 40.5% increase in completed PRs relative to zero-usage weeks, with a monotonic gradient across Low, Moderate, and High usage and evidence of diminishing returns at the top end. They then run seven falsification checks, including a non-coding Microsoft 365 Copilot placebo, a teammate PR placebo outcome, lag/lead timing tests, PR-size decomposition, PR-type decomposition, and an alternate breadth-based treatment definition; these largely rule out generic AI engagement, team-level shocks, within-week task reallocation, slicing into smaller PRs, and easier task mix as simple explanations. The breadth-based robustness check still shows a 54% increase at high usage, reinforcing the same qualitative conclusion: within-engineer Copilot-heavy weeks are associated with substantially higher PR throughput, though the paper emphasizes that this remains observational and likely understates any total effect if Copilot also saves time.

Original abstract

Does GitHub Copilot (GHCP) make engineers more productive, or do the engineers who use it more differ from those who use it less? And even within a single engineer, are GHCP-heavy weeks just busy weeks in which more of everything gets done? We study these questions using 43 weeks of data from 16,223 software engineers across Microsoft's Cloud+AI organization. Engineer fixed effects address the first concern by comparing each engineer against themselves rather than against other engineers, eliminating time-invariant differences in skill, role, and team. Active coding time and browser time then enter a Poisson Pseudo-Maximum Likelihood model with two-way fixed effects to address the harder, within-engineer confound: that GHCP-heavy weeks coincide with high-effort weeks. This defines our estimand as an efficiency effect: more pull requests completed at equivalent levels of coding time. Engineers are estimated to complete 40.5% more PRs in their highest GHCP usage weeks relative to their zero-usage weeks, holding measured development effort constant. The gradient is monotonic with diminishing returns at high intensity. Seven robustness and falsification tests target the remaining plausible alternative explanations (non-coding AI engagement, team-level shocks, within-week task reallocation, cross-week contamination, PR slicing into smaller units, shifts toward easier task types, and sensitivity to how the treatment is operationalized). Under an explicitly stated conditional-independence assumption, the within-engineer design estimates a tool-specific efficiency effect that is consistent with all seven robustness tests.

Read the original paper

More in Code Generation

Browse all 43 papers →
02Code Generation

Groupwise Agentic Grading and Advantage Redistribution for Code Agent RL

Jinhao Dong, Liang Zhao, Zihao Yue, Wenhan Ma, Linghao Zhang, Lei Li, Shicheng Li, Yifan Song, Bowen Ye, Fuli Luo

GAGAR helps code agents learn not only to pass tests, but to produce cleaner and more targeted implementations by redistributing RL credit according to agentic quality judgments.

Read analysis