NTH
AI research

RoadmapBench: Evaluating Long-Horizon Agentic Software Development Across Version Upgrades

AuthorsXinbo Xu, Ruihan Yang, Haiyang Shen, Wendong Xu, Bofei Gao, Ruoyu Wu, Kean Shi, Weichu Xie, Xuanzhong Chen, Ming Wu, Jason Zeng, Michael Heinrich, Elvis Zhang, Liang Chen, Kuan Li, Baobao Chang

May 18, 2026 2 min read
Watch on YouTube
The one-line take

RoadmapBench is a new benchmark showing that today’s top coding agents still struggle badly with realistic, large-scale software version upgrades spanning many files and months of work.

Key results

115
tasks

RoadmapBench contains 115 long-horizon coding tasks built from real open-source version upgrades.

17
repositories

The benchmark spans 17 open-source repositories to cover diverse software domains.

5
languages

Tasks cover 5 programming languages, broadening beyond Python-only coding benchmarks.

3,714 lines changed across 51 files
median oracle patch

A typical ground-truth upgrade requires multi-file, multi-module edits at real engineering scale.

about 5 subtasks
average subtasks per task

Each task is decomposed into multiple roadmap targets with weighted subtask scoring.

Claude-Opus-4.7: 39.1% resolved, 0.692 Completion Score
best model result

Under OpenHands, the strongest evaluated model solves only a minority of tasks, while the Completion Score captures partial progress.

What the paper found

RoadmapBench introduces a new benchmark for long-horizon agentic software development by converting real open-source version upgrades into 115 multi-target coding tasks across 17 repositories and 5 languages, with a median oracle patch of 3,714 lines changed across 51 files and about 5 subtasks per task. Unlike prior bug-fix benchmarks such as SWE-bench Verified, each instance uses a roadmap-style specification that exposes behavioral targets while hiding implementation details, and each subtask has its own weighted test score so partial progress is measurable via Completion Score instead of binary pass/fail alone. Evaluating 13 frontier models under OpenHands, the strongest system, Claude-Opus-4.7, resolves only 39.1% of tasks with a 0.692 Completion Score, while Seed-2.0-Pro reaches 5.2%, showing that multi-file version upgrades remain largely unsolved. The benchmark also diagnoses failure structure: performance drops monotonically with more files, more changed lines, and higher subtask counts; strong models fail mostly through implementation-level bugs, while weaker models fail earlier with build errors, missing implementations, and interface mismatches. The authors also validate task quality with static checks plus rollout-based attribution, repairing confirmed task-side defects before release, which makes RoadmapBench notable not just as a harder benchmark, but as a finer-grained diagnostic of where coding agents break down during realistic software evolution.

Original abstract

Coding agents are increasingly deployed in real software development, where a single version iteration requires months of coordinated work across many files. However, most existing benchmarks focus predominantly on single-issue bug fixes from Python repositories, with coarse pass/fail evaluation outcomes, and thus fail to capture long-horizon, multi-target development at real engineering scale. To address this gap, we present RoadmapBench, a benchmark of 115 long-horizon coding tasks grounded in real open-source version upgrades across 17 repositories and 5 programming languages. Each task places the agent on a source-version code snapshot and provides a multi-target roadmap instruction requiring it to implement the functionality introduced in the target version, with a median modification of 3,700 lines across 51 files. We conduct a systematic evaluation on thirteen frontier models and find that even the strongest, Claude-Opus-4.7, resolves only 39.1% of tasks, while the weakest achieves merely 5.2%, in stark contrast to existing bug-fix benchmarks, suggesting that long-horizon software development remains a largely unsolved problem.

Read the original paper