NTH

SWE-Review: Closing the Loop on Issue Resolution with Agentic Code Review

AuthorsRuoyu Wang, Jierun Chen, Shaowei Wang, Chaofan Tao, Sidi Yang, Yuxin Jiang, Kim-Hui Yap, Lifeng Shang, Xiaohui Li, Haoli Bai

July 14, 2026 2 min read
Watch on YouTube
The one-line take

This paper teaches coding agents not just to write pull requests, but to review, critique, and revise them in a loop that more closely matches real software development.

Key results

1384
SWE-Review-Bench PRs

candidate pull requests in the benchmark

500
SWE-bench Verified issues

issues underlying SWE-Review-Bench

8914
SWE-Review-Traj trajectories

decision-correct review trajectories

27.5%
Qwen3-30B-A3B RRR baseline

no-review resolve rate on SWE-bench Verified

52.6%
Qwen3-30B-A3B RRR with agentic review

resolve rate after revision with agentic review

What the paper found

SWE-Review from NTU, Huawei Technologies, and HKU reframes AI code review as a closed-loop mechanism for issue resolution rather than a post-hoc comment generator. The paper builds SWE-Review-Bench with 1,384 AI-generated pull requests from 500 SWE-bench Verified issues and SWE-Review-Traj with 8,914 decision-correct review trajectories, then evaluates agentic reviewers such as Claude Opus 4.6, GLM-5, and Qwen3 variants. Compared with single-turn fixed-context review, agentic review improves both decision accuracy and downstream resolve rate, especially on non-local bugs: for Qwen3-30B-A3B, resolve rate after revision rises from 27.5% to 52.6%, while the generate-review-revise loop lifts Qwen3-30B-A3B from 27.5% to 56.9% and Qwen3-Coder-30B-A3B from 50.9% to 68.8%. The strongest reviewer, Claude Opus 4.6, reaches 89.4% decision accuracy on one split and 81.8% weighted-average accuracy with only 148K tokens on average, and review-guided iterative revision outperforms verifier-selected best-of-N by reaching 38.4% resolve rate from a 22.9% baseline using only 2.44 samples on average. The key technical claim is that structured, repository-exploring review feedback is not just a gate, but a reusable supervision signal for training, revision, and test-time scaling.

Original abstract

Coding agents increasingly generate pull requests (PRs) for real-world software issues, yet one-shot PR generation remains open-loop: the PR is proposed without systematic review, diagnosis, or revision. We introduce \textbf{SWE-Review}, a framework for closing this loop with agentic code review. Given an issue and an AI-generated PR, a reviewer agent explores the repository, decides whether the PR should be accepted, and provides structured feedback for revision. We evaluate this setting with our proposed \textbf{SWE-Review-Bench} to measure both review correctness and downstream revision usefulness. We further curate \textbf{SWE-Review-Traj} dataset to study broader applications of agentic review and fill the data-scarcity gap for open reviewer training. Experiments show that agentic review continuously improves PRs through a generate-review-revise loop, outperforms single-turn fixed-context review in both decision accuracy and resolve rate after revision, transfers beyond review to improve issue-resolution models, and enables effective and efficient test-time scaling. These results position agentic code review as a practical mechanism for moving AI coding agents from one-shot PR generation toward closed-loop issue resolution.

Read the original paper

More in Code Generation

Browse all 43 papers →
02Code Generation

Groupwise Agentic Grading and Advantage Redistribution for Code Agent RL

Jinhao Dong, Liang Zhao, Zihao Yue, Wenhan Ma, Linghao Zhang, Lei Li, Shicheng Li, Yifan Song, Bowen Ye, Fuli Luo

GAGAR helps code agents learn not only to pass tests, but to produce cleaner and more targeted implementations by redistributing RL credit according to agentic quality judgments.

Read analysis