From Human-Centric to Agentic Code Review: The Impact of Different Generations of Generative AI Technology on Review Quality
AuthorsSuzhen Zhong, Shayan Noei, Bram Adams, Ying Zou
Resources
A large study finds that AI agents can speed up code reviews, but faster decisions do not necessarily mean higher-quality reviews.
Key results
Longitudinal review dataset spanning the transition across three AI review eras.
Open-source projects with sufficient review activity in all three eras.
Days/KLOC reduction from the pre-LLM era to the agent era.
Days/KLOC reduction from the pre-LLM era to the agent era.
Increase in review-smell prevalence during the LLM era.
Average prevalence of repeated-reviewer reliance in LLM-involved collaboration patterns.
What the paper found
This study by Suzhen Zhong, Shayan Noei, Bram Adams, and Ying Zou examines how code review changes as projects move from human-centric review to LLM-assisted and agentic review. Using 1.02M pull requests across 207 GitHub projects, the authors define three adoption trajectories—Gradual AI Adoption, Rapid LLM Adoption, and Rapid AI Agent Adoption—and analyze them with soft-DTW clustering, Markov-chain expectation maximization, and logistic regression. Agentic reviewers, including tools such as Anthropic’s Claude Code, can retrieve context, run tools, and verify findings rather than merely generate comments. Gradual AI Adoption reduced review delay by 2.5 days/KLOC in the agent era, while Rapid AI Agent Adoption reduced it by 4.5 days/KLOC; however, Rapid LLM Adoption increased review-smell prevalence by 8.0% and produced no significant efficiency gain. Across ten human-AI collaboration patterns, agent-initiated and multi-agent reviews were faster than human-only review under gradual or rapid-agent adoption, but AI-involved reviews generally carried more quality risk. The dominant risk was Review Buddies: repeated reliance on the same reviewer rose to 60% for LLM-involved patterns, compared with 16% for human-only review. The analysis also finds that pull-request type, changeset size, review activity, and author experience remain important, while collaboration patterns become the strongest efficiency-related factor after AI reviewers enter. Projects associated with Microsoft and Google illustrate the rapid-agent adoption trajectory, but the authors caution that these are explanatory associations, not causal effects, and recommend context-aware, selective AI deployment rather than uniform automation.
Original abstract
Code review helps maintain software quality before code integration, but it also imposes a substantial workload on human reviewers. As generative artificial intelligence becomes part of software development, code review is shifting from a primarily human review process toward AI-supported review processes in which large language model (LLM) reviewers and AI agent reviewers participate alongside human reviewers. However, we still lack empirical evidence on how this transition affects review efficiency and review quality. In this paper, we study 1.02 million reviewed pull requests from 207 GitHub projects that transition across three code review eras: human-centric review, LLM-assisted review, and agentic code review. We identify three AI reviewer adoption practices: Gradual AI Adoption, Rapid LLM Adoption, and Rapid AI Agent Adoption. We further model pull request review discussions as reviewer interaction sequences to characterize how human, LLM, and AI agent reviewers collaborate during the review process. Our results show that agent-involved collaboration patterns, especially reviews initiated by AI agents or involving multiple AI agents, are associated with faster review decisions under Gradual AI Adoption and Rapid AI Agent Adoption. However, these efficiency gains do not translate into better review quality. We also find that review activity and pull request type remain important across eras, while human-AI collaboration patterns become the strongest explanatory factor for review efficiency once LLM and AI agent reviewers participate. These findings provide empirical guidance for designing AI-supported code review processes that improve efficiency without weakening review quality.
Read the original paperMore in Code Generation
Browse all 43 papers →Compact Documentation for Coding Agents: A Benchmark, an Optimizer, and Why It Does Not Transfer
Md Shohel Arman, Igor Molybog
Better code documentation can faithfully reconstruct software, but surprisingly does not necessarily help AI coding agents fix real issues when the source code is already available.
Groupwise Agentic Grading and Advantage Redistribution for Code Agent RL
Jinhao Dong, Liang Zhao, Zihao Yue, Wenhan Ma, Linghao Zhang, Lei Li, Shicheng Li, Yifan Song, Bowen Ye, Fuli Luo
GAGAR helps code agents learn not only to pass tests, but to produce cleaner and more targeted implementations by redistributing RL credit according to agentic quality judgments.
Reinforcement Learning from Intermediate Renders for Image-to-Code Generation
Omri Kaduri, Kate Feingold, Phillip Isola, Tali Dekel
IR4RL improves image-to-code generation by rewarding models for making useful visual progress at every intermediate rendering step.