NTH

SWE Refactor Bench: Can Coding Agents Complete a Long-Horizon, Whole-Repository Stack Migration?

AuthorsDeyao Hong, Yizhe Chi, Wenyi Li, Xiaoqiu Wang, Mingju Gao, Kaisen Yang, Bingxiang He, Youjie Zheng, Calvin Xiao, Qinhuai Na

August 30, 2026 3 min read
Watch on YouTube
The one-line take

SWE Refactor Bench tests whether coding agents can truly migrate an entire codebase without merely making the tests pass, and finds that they still struggle significantly.

Key results

130,118
Fixed behavioral checks

Differential checks used in the Behavioural Tests stage

520
Scored runs

Runs across eight frontier models and 26 effort configurations

5.4%
Accepted migrations

28 of 520 runs passed Migration Audit, Behavioural Tests, and Agentic Verification

47.0
Best model score

Claude Opus 5 at xhigh effort, out of 100

What the paper found

SWE Refactor Bench evaluates whether coding agents can perform long-horizon, whole-repository stack migrations rather than merely make existing tests pass. Its 20 migrations cover language, framework, platform, and build-toolchain debt in projects including SQLite, zlib, libsodium, and GraphHopper. The benchmark addresses “Blindness,” where behavior-only evaluation rewards an untouched repository, with three sequential gates: Migration Audit confirms that the old stack was removed, Behavioural Tests enforce 130,118 fixed differential checks, and Agentic Verification uses six independent coding agents for one hour each to generate executable counterexamples. Across 520 runs spanning eight frontier models and 26 effort configurations—including Claude Opus 5, Claude Sonnet 5, GPT-5.6, Kimi K3, Qwen 3.8 Max, and DeepSeek V4 Flash—only 28 runs, or 5.4%, passed all three stages, while 13 tasks received no accepted solution. Claude Opus 5 at xhigh effort led with 47.0 out of 100. The failures separate migration from correctness: 30 runs preserved behavior by skipping the migration, whereas 252 migrated but broke behavior. Even among migrations that passed the audit, only 26% passed every fixed check, and agentic verification rejected 60 of 88 submissions that reached it. Capability varied sharply by category: build-toolchain rewrites averaged 31.4, while language rewrites averaged just 5.6. The result is a benchmark showing that repository migration requires both structural replacement and exact behavioral preservation, with hidden differential testing exposing regressions that fixed suites miss.

Original abstract

Modern software systems accumulate technical debt over decades of development, which makes migration expensive and largely manual. As coding agents become increasingly capable at bug fixing, can they autonomously perform such migrations? Existing benchmarks cannot answer this question because they evaluate only behavioural correctness, not whether the migration actually occurred. This leads an easy hack: agents copy the original implementation to make tests pass. We call this Blindness. To address this problem, we introduce SWE Refactor Bench, a benchmark comprising 20 whole-repository migrations, covering 4 kinds of technical debt. A three-stage evaluation protocol measures both migration completeness and behavioural correctness. (1) Migration Audit verifies that the migration occurred. (2) Behavioural Tests measure correctness with a fixed test suite. (3) Agentic Verification uses 6 independent coding agents to generate targeted tests for hidden behavioural differences. Across 520 runs from 8 frontier models and 26 model-effort configurations, only 28 of 520 runs ($5.4\%$) pass all three stages, 13 of the 20 tasks receive no accepted solution, and the best model (claude-opus-5) scores $47.0/100$. Migration completeness and behavioural correctness are distinct abilities: a few runs preserve behaviour by skipping the migration and are stopped at Migration Audit; most attempt it and break behaviour, and are stopped at Behavioural Tests. Agents cannot deliver a perfect migration: among the 340 runs that pass Migration Audit, $58\%$ reach $99\%$ of the fixed checks, yet only $26\%$ reach $100\%$. Agent capability differs across migration categories: agents score $31.4$ on build toolchain rewrites but only $5.6$ on language rewrites. Together, these findings position SWE Refactor Bench as a rigorous testbed for developing coding agents for reliable whole-repository migrations.

Read the original paper

More in AI Benchmarks

Browse all 45 papers →
01Benchmark

Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows

Gabriel Tomitsuka, Arman Raayatsanati, Emma Xing, Duke Gand, Joseph J Ma

Argo-Bench stress-tests AI data agents on realistic enterprise-scale workflows where success depends not just on writing SQL, but on correctly investigating data and taking actions with real consequences.

Read analysis
02Benchmark

EnigmaForge: The Question Is Hidden in the Story

Daniel Eisner

EnigmaForge tests whether AI can discover and solve a hidden puzzle in a story, revealing reasoning abilities that ordinary question-answering benchmarks may miss.

Read analysis