NTH

To Run or Not to Run: Analyzing the Cost-Effectiveness of Code Execution in LLM-Based Program Repair

AuthorsZhihao Lin, Junhua Zhu, Mingyi Zhou, Xin Wang, Zhensu Sun, Renyu Yang, David Lo, Li Li

July 5, 2026 3 min read
Watch on YouTube
The one-line take

This paper asks whether running tests during LLM code repair is really worth it, and finds that agents often pay a big execution cost for only modest gains.

Key results

7745
agent traces analyzed

public SWE-bench leaderboard traces used for execution-behavior analysis

3000
controlled repair attempts

end-to-end repairs across 200 SWE-bench instances, 3 agents, and 5 execution modes

200
SWE-bench instances

first 100 from SWE-bench Lite and first 100 from SWE-bench Verified

8.8
average test runs per task

mean execution frequency across analyzed traces

1.25
resolve-rate gap

average Prohibited minus Unrestricted gap in percentage points for commercial agents

56-62%
token savings

Claude Code Prohibited versus Unrestricted token reduction

What the paper found

This ISSTA 2026 paper asks whether iterative code execution is actually worth its cost in LLM-based program repair, using Claude Code from Anthropic, OpenAI Codex, and the open-source OpenCode with Qwen2.5-Coder-32B. Across 7,745 SWE-bench leaderboard traces and 3,000 controlled repair attempts on 200 SWE-bench instances, the authors isolate execution access as the only variable across four paradigms. They find that agents execute frequently—8.8 test runs per task on average, ranging from 2 to 19—but the marginal value is small: the Prohibited versus Unrestricted resolve-rate gap averages only 1.25 percentage points for commercial agents and is not statistically significant. On Claude Code, Prohibited reduces token use by 56–62% and wall-clock time by 48–54% while keeping resolve rate within 1–3 points of Unrestricted; on OpenCode, the gap is approximately 0 points. The study also shows that execution benefit is highly uneven: 54–66% of commercial-agent cases finish in a single edit, 48.8% of Claude Code’s reproduction executions produce actionable localization feedback, and 81–100% of failed cases can pass the agent’s own validation but still fail SWE-bench’s official tests. Late-stage executions perform better than early ones, with an average success rate of 57.9%, but overall the paper concludes that code execution should be treated as an explicit resource with a cost-benefit threshold, not a default capability.

Original abstract

LLM-based agents for program repair are increasingly built on a "generate-run-revise" paradigm, iteratively executing tests to evaluate and refine patches. This execution-based approach has become standard practice in state-of-the-art systems. However, executions can be time-consuming and expensive, yet their impact on these agents remains underexplored. In this paper, we conduct a two-stage empirical study over execution behavior in LLM-based program repair. To characterize execution behavior at scale, we first analyze 7,745 agent traces from SWE-bench leaderboard submissions. Second, we evaluate 3,000 end-to-end repair attempts across 200 SWE-bench instances and three agents (Claude Code, Codex, and the open-source OpenCode) under four execution paradigms, which allows for a fine-grained comparison of performance and cost. Our analysis reveals three key observations: (1) Code execution is used across all agents and models analyzed, with an average of 8.8 test runs per task. Execution behavior varies substantially across agents and models, with frequency ranging from 2 to 19 per task, and late-stage executions consistently achieve higher success rates than early-stage ones. (2) Execution restrictions have little effect on repair success: on commercial agents with SOTA models the resolve-rate gap between Prohibited and Unrestricted is only 1.25 percentage points and not statistically significant, while Prohibited saves substantial token and wall-clock cost. (3) Execution benefit is concentrated rather than uniform. These patterns suggest that current agents apply execution indiscriminately, paying its cost on instances where it provides little benefit. Execution, therefore, should be treated as a resource with an explicit cost-benefit tradeoff, not a default capability.

Read the original paper

More in Code Generation

Browse all 43 papers →
02Code Generation

Groupwise Agentic Grading and Advantage Redistribution for Code Agent RL

Jinhao Dong, Liang Zhao, Zihao Yue, Wenhan Ma, Linghao Zhang, Lei Li, Shicheng Li, Yifan Song, Bowen Ye, Fuli Luo

GAGAR helps code agents learn not only to pass tests, but to produce cleaner and more targeted implementations by redistributing RL credit according to agentic quality judgments.

Read analysis