To Run or Not to Run: Analyzing the Cost-Effectiveness of Code Execution in LLM-Based Program Repair
AuthorsZhihao Lin, Junhua Zhu, Mingyi Zhou, Xin Wang, Zhensu Sun, Renyu Yang, David Lo, Li Li
Resources
This paper asks whether running tests during LLM code repair is really worth it, and finds that agents often pay a big execution cost for only modest gains.
Key results
public SWE-bench leaderboard traces used for execution-behavior analysis
end-to-end repairs across 200 SWE-bench instances, 3 agents, and 5 execution modes
first 100 from SWE-bench Lite and first 100 from SWE-bench Verified
mean execution frequency across analyzed traces
average Prohibited minus Unrestricted gap in percentage points for commercial agents
Claude Code Prohibited versus Unrestricted token reduction
What the paper found
This ISSTA 2026 paper asks whether iterative code execution is actually worth its cost in LLM-based program repair, using Claude Code from Anthropic, OpenAI Codex, and the open-source OpenCode with Qwen2.5-Coder-32B. Across 7,745 SWE-bench leaderboard traces and 3,000 controlled repair attempts on 200 SWE-bench instances, the authors isolate execution access as the only variable across four paradigms. They find that agents execute frequently—8.8 test runs per task on average, ranging from 2 to 19—but the marginal value is small: the Prohibited versus Unrestricted resolve-rate gap averages only 1.25 percentage points for commercial agents and is not statistically significant. On Claude Code, Prohibited reduces token use by 56–62% and wall-clock time by 48–54% while keeping resolve rate within 1–3 points of Unrestricted; on OpenCode, the gap is approximately 0 points. The study also shows that execution benefit is highly uneven: 54–66% of commercial-agent cases finish in a single edit, 48.8% of Claude Code’s reproduction executions produce actionable localization feedback, and 81–100% of failed cases can pass the agent’s own validation but still fail SWE-bench’s official tests. Late-stage executions perform better than early ones, with an average success rate of 57.9%, but overall the paper concludes that code execution should be treated as an explicit resource with a cost-benefit threshold, not a default capability.
Original abstract
LLM-based agents for program repair are increasingly built on a "generate-run-revise" paradigm, iteratively executing tests to evaluate and refine patches. This execution-based approach has become standard practice in state-of-the-art systems. However, executions can be time-consuming and expensive, yet their impact on these agents remains underexplored. In this paper, we conduct a two-stage empirical study over execution behavior in LLM-based program repair. To characterize execution behavior at scale, we first analyze 7,745 agent traces from SWE-bench leaderboard submissions. Second, we evaluate 3,000 end-to-end repair attempts across 200 SWE-bench instances and three agents (Claude Code, Codex, and the open-source OpenCode) under four execution paradigms, which allows for a fine-grained comparison of performance and cost. Our analysis reveals three key observations: (1) Code execution is used across all agents and models analyzed, with an average of 8.8 test runs per task. Execution behavior varies substantially across agents and models, with frequency ranging from 2 to 19 per task, and late-stage executions consistently achieve higher success rates than early-stage ones. (2) Execution restrictions have little effect on repair success: on commercial agents with SOTA models the resolve-rate gap between Prohibited and Unrestricted is only 1.25 percentage points and not statistically significant, while Prohibited saves substantial token and wall-clock cost. (3) Execution benefit is concentrated rather than uniform. These patterns suggest that current agents apply execution indiscriminately, paying its cost on instances where it provides little benefit. Execution, therefore, should be treated as a resource with an explicit cost-benefit tradeoff, not a default capability.
Read the original paperMore in Code Generation
Browse all 43 papers →Compact Documentation for Coding Agents: A Benchmark, an Optimizer, and Why It Does Not Transfer
Md Shohel Arman, Igor Molybog
Better code documentation can faithfully reconstruct software, but surprisingly does not necessarily help AI coding agents fix real issues when the source code is already available.
Groupwise Agentic Grading and Advantage Redistribution for Code Agent RL
Jinhao Dong, Liang Zhao, Zihao Yue, Wenhan Ma, Linghao Zhang, Lei Li, Shicheng Li, Yifan Song, Bowen Ye, Fuli Luo
GAGAR helps code agents learn not only to pass tests, but to produce cleaner and more targeted implementations by redistributing RL credit according to agentic quality judgments.
Reinforcement Learning from Intermediate Renders for Image-to-Code Generation
Omri Kaduri, Kate Feingold, Phillip Isola, Tali Dekel
IR4RL improves image-to-code generation by rewarding models for making useful visual progress at every intermediate rendering step.