NTH

Reusing Past Repairs Through Hierarchical Trajectory Abstraction for Coding Agents

AuthorsYisen Xu, Jiayuan Zhou, Ruiqi Pan, Tse-Hsun Chen

August 4, 2026 2 min read
Watch on YouTube
The one-line take

STAIR helps coding agents learn from past software repairs by turning old debugging trajectories into reusable, adaptable plans.

Key results

81.2%
STAIR with MiniMax M2.5

Pass@1 on SWE-bench Verified

79.2%
STAIR with GPT-5

Pass@1 on SWE-bench Verified

500
Benchmark size

Instances in SWE-bench Verified

81.0%
Cross-agent transfer

Pass@1 for mini-SWE-agent v2 with STAIR plans

57.6%
Raw-trajectory ablation

Resolution rate on the 125-instance ablation subset

22.4%
Hierarchy benefit

Performance drop from full hierarchical abstraction to raw trajectories

What the paper found

Researchers from Concordia University’s SPEAR Lab and Huawei Canada introduce STAIR, a framework that lets coding agents reuse procedural knowledge from earlier repairs instead of treating every issue independently. STAIR divides successful Lingxi repair trajectories into localization, planning, and execution-and-verification stages, then uses GPT-5 to organize each stage into a hierarchy ranging from repository-specific actions to transferable debugging strategies. For a new issue, similarity retrieval and an LLM relevance verifier select nodes across abstraction levels, while an LLM adapts them into executable, stage-specific plans. On the 500-instance SWE-bench Verified benchmark, STAIR reaches 81.2% Pass@1 with MiniMax M2.5 and 79.2% with GPT-5, outperforming same-backbone baselines including mini-SWE-agent v2, which scores 75.8% with MiniMax M2.5. The plans also transfer without code changes to mini-SWE-agent v2, raising Pass@1 to 81.0%, an absolute gain of 5.2%, while reducing average token use from 1,245k to 1,180k per instance. Ablations show why hierarchy matters: on a difficulty-stratified 125-instance subset, the full system resolves 80.0%, compared with 57.6% for raw trajectories, a 22.4% drop. The results indicate that reusable repair knowledge must combine concrete actions, intermediate investigation strategies, and high-level principles, then be adapted to the target repository rather than pasted into an agent prompt as generic advice.

Original abstract

Although LLM-driven repair agents can tackle complex, repository-level issues, they treat every issue independently and discard the procedural knowledge accumulated from previous repairs. We introduce STAIR, a framework that converts historical repair trajectories into hierarchical, reusable plans that can be adapted to steer future repairs. Each past trajectory is transformed into a multi-level tree that ranges from fine-grained diagnostic actions to high-level repair strategies, encoding experience at several granularities. When a new issue arrives, STAIR selects relevant plan nodes from multiple abstraction levels, tailors them into executable, issue-specific plans, and supplies them to the agent through its prompt. On SWE-bench Verified, STAIR integrated with Lingxi reaches 81.2% Pass@1 using MiniMax M2.5 and 79.2% using GPT-5. The generated plans also generalize across agents: without any code change, they lift the Pass@1 of a structurally different agent, mini-SWE-agent v2, from 75.8% to 81.0%. Ablation experiments further show that mixing multiple abstraction levels surpasses any single level and that raw, unabstracted trajectories transfer substantially worse.

Read the original paper

More in AI Agents

Browse all 56 papers →
01Agent

LEGO-Anything: Coding Agents for 3D Scene Reconstruction

Xirui Li, Peng Shi, Mingwen Dong, Sheng Zhang, Zhuoyan Xu, Dongkyu Lee, Shuaichen Chang, Yi Xiang, Lin Pan, Jiarong Jiang

LEGO-Anything turns images into editable Blender programs through iterative coding agents, offering a promising but still imperfect route to reconstructable 3D worlds.

Read analysis
02Agent

MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution

Prithwish Jana, Mononito Goswami, Hao Liu, Xinyu Li, Langlin Huang, Zhehui Huang, Zhishen Huang, Patrick Blöbaum, Anoop Deoras, Purak Jain, Nikos Kanakaris, Sahika Genc

MILO uses teams of evolving AI agents to automatically discover better harnesses for long-horizon problem-solving systems.

Read analysis
03Agent

Self-Organizing Agent Teams Learn to Reason Together

Aneesh Pappu, Mirac Suzgun, Yongchan Kwon, Federico Bianchi, Batu El, Mykel J. Kochenderfer, Hancheng Cao, James Zou

This work trains AI agents to discover how to divide labor, challenge ideas, and combine reasoning so that teams can solve problems no individual agent could solve alone.

Read analysis