ExecCritic: Learn to Test, Test to Improve for Coding Agents
AuthorsLeitian Tao, Baolin Peng, Haorui Wang, Hang Wang, Hao Cheng, Wenlin Yao, Qianhui Wu, Tao Ge, Sharon Li, Jianfeng Gao
ExecCritic trains one coding agent to create trustworthy tests and another to repair code from them, substantially improving repository-level bug fixing.
Key results
Resolved rate for the jointly post-trained Qwen Test and Repair agents.
Percentage-point improvement from composing the trained agents.
Resolved rate versus the 61.2% no-test baseline for a fixed Repair agent.
Resolved rate achieved by the same fixed Repair agent using OpenAI GPT-5.6-sol-generated tests.
Post-trained Qwen Test-agent success, up from 22.2% before post-training.
Additional agent turns among 120 trajectories that entered feedback-guided repair.
What the paper found
ExecCritic addresses a failure mode in coding agents: a patch and its self-generated test can share the same mistaken interpretation, producing false confidence. Its test–verify–revise scaffold assigns separate roles to a Test agent and a Repair agent, using Qwen-3.5-35B-A3B for both, while a fail-closed harness validates a repository-native test, freezes it, and exposes only bounded execution feedback during source-only repair. Learn to Test uses supervised fine-tuning on 5,000 SWE-ReBench trajectories from DeepSeek-V4-Flash-0731, followed by GRPO rewarding Base-to-Gold validity and discrimination between correct and incorrect patches; Test to Improve trains repair from fixed feedback. On SWE-bench Verified, Qwen-generated tests lower a fixed Repair agent’s resolved rate from 61.2% without tests to 57.3%, while GPT-5.6-sol tests from OpenAI raise it to 65.3%, showing that feedback quality is decisive. Post-training lifts Qwen test reliability from 22.2% to 62.2% Base-to-Gold success, and composing the trained Test and Repair agents reaches 72.6% resolution—an 11.4-point gain over the original no-test baseline—without GPT-5.6-sol or Oracle feedback at evaluation. Microsoft’s Orchard framework supports the scalable agent training and sandbox execution. Among 120 trajectories entering repair, revision adds an average of 13 agent turns, indicating measurable gains without routinely exhausting the five-round revision budget.
Original abstract
Execution feedback can guide coding agents toward correct repository repairs, but only when the tests capture the behavior requested by the issue. Agent-generated tests can encode incomplete or incorrect behavioral targets; when the same trajectory writes both the patch and the test, their errors can agree and create false confidence. We introduce ExecCritic, combining a test--verify--revise scaffold with a role-specific reinforcement learning recipe for training agents within it. The scaffold separates test construction from source-code repair: a Test agent independently generates repository-native tests, a fail-closed harness qualifies and freezes them, and a Repair agent revises source code from their execution feedback without changing the tests. Both roles use Qwen-3.5-35B-A3B as the backbone and are trained separately. In Learn to Test, the Test agent learns to produce behaviorally valid tests that distinguish correct from incorrect patches. In Test to Improve, the Repair agent learns both direct task resolution and feedback-guided revision. On SWE-bench Verified, test quality determines whether feedback helps: holding the base Repair agent fixed, tests from the base Test agent reduce resolved rate from a no-test baseline of 61.2% to 57.3%, whereas tests from GPT-5.6-sol raise it to 65.3%. Role-specific post-training raises the Qwen Test agent's Base-to-Gold success from 22.2% to 62.2%; composing the two post-trained Qwen agents reaches 72.6%, an 11.4-point gain over the original no-test baseline without stronger-model or Oracle feedback at evaluation time. Code is publicly available at https://github.com/MSR-Orchard/execcritic.
Read the original paperMore in Code Generation
Browse all 43 papers →Compact Documentation for Coding Agents: A Benchmark, an Optimizer, and Why It Does Not Transfer
Md Shohel Arman, Igor Molybog
Better code documentation can faithfully reconstruct software, but surprisingly does not necessarily help AI coding agents fix real issues when the source code is already available.
Groupwise Agentic Grading and Advantage Redistribution for Code Agent RL
Jinhao Dong, Liang Zhao, Zihao Yue, Wenhan Ma, Linghao Zhang, Lei Li, Shicheng Li, Yifan Song, Bowen Ye, Fuli Luo
GAGAR helps code agents learn not only to pass tests, but to produce cleaner and more targeted implementations by redistributing RL credit according to agentic quality judgments.
Reinforcement Learning from Intermediate Renders for Image-to-Code Generation
Omri Kaduri, Kate Feingold, Phillip Isola, Tali Dekel
IR4RL improves image-to-code generation by rewarding models for making useful visual progress at every intermediate rendering step.