NTH

Change2Task: From Repository Changes to Executable Coding Agent Tasks and Environments

AuthorsHaomin Qi, Xingliang Wang, Xuanqi Gao, Baihui Sang, Xin Zhang, Minghua Ma, Pengfei Gao, Yu Kang, Qingwei Lin, Saravan Rajmohan, Dongmei Zhang, Qi Zhang

August 5, 2026 2 min read
Watch on YouTube
The one-line take

Change2Task turns real pull-request history into reliable, reusable coding-agent tasks for training and evaluation.

Key results

1130
Eligible source changes

Repository changes eligible for task construction.

79.6%
Finalized task recovery

900 verified tasks recovered from 1130 eligible changes.

900
Finalized task pairs

Verified task pairs spanning five coding-agent task families.

29.2%
Bug Fix recovery improvement

Relative improvement over SWE-smith PR Mirror on matched candidates.

0.894
Source-change fidelity

Task-weighted fidelity between historical and reconstructed changes.

58.4%
Environment setup-time reduction

Reduction from reusing modern repository bases.

What the paper found

Microsoft researchers introduce Change2Task, a framework that converts merged pull requests into executable coding-agent tasks on healthy, modern descendant revisions of the same repository. Instead of replaying historical code directly, it preserves developer intent through three escalating methods—Patch Reversal, Code Mapping, and Agent Reconstruction—then validates a complete healthy-to-task-to-restored lifecycle with target checks, regression checks, scope constraints, and source-change fidelity. Across five task families—Bug Fix, Feature Addition, Test Generation, API Migration, and Security Repair—Change2Task processes 1130 eligible changes and finalizes 900 verified task pairs, for a 79.6% recovery rate. On 621 matched Bug Fix candidates, it recovers 500 tasks versus 387 for SWE-smith PR Mirror, a 29.2% relative improvement. The resulting corpus reaches 0.894 task-weighted source-change profile fidelity, while paired evaluations using Codex with GPT-5.5, Claude Code with Sonnet 5, Gemini CLI with Gemini 3.1 Pro, and GitHub Copilot with GPT-5.6 Terra show that reconstructed tasks preserve historical evaluation signals. Reusing 388 modern bases cuts environment setup time by 58.4% and retained storage by 71.2%, demonstrating that one executable repository environment can support multiple provenance-linked tasks. The system therefore turns repository history into scalable, verifiable infrastructure for coding-agent training and benchmarking, complementing tools such as SWE-bench, SWE-smith, and RepoLaunch.

Original abstract

Scaling coding agents requires a continuing supply of executable data for training, benchmarking, and continuous evaluation. Each task must couple a realistic software state with a specification, development tools, and reliable verification. To expand this supply, we present Change2Task, a system grounded in repository history that converts merged pull requests into verified tasks on healthy modern revisions of the same repository. It aligns historical evidence with evolved code, reconstructs task states through Patch Reversal, Code Mapping, or Agent Reconstruction, and validates the lifecycle from a healthy base to a task state and a restored state. By deriving multiple tasks grounded in developer evidence from maintained environments, Change2Task provides executable data for coding agent training and evaluation while reducing repeated environment setup, storage, and task construction effort. We evaluate the system through five common and widely adopted coding agent task families: Bug Fix, Feature Addition, Test Generation, Application Programming Interface Migration, and Security Repair. Starting from 1,130 source changes eligible for construction, Change2Task achieves 79.6% verified task construction success across these task families. On a matched candidate set, it recovers 29.2% more verified tasks than a construction baseline based on pull requests. Historical and reconstructed cases achieve up to 98.0% matched outcome agreement under agent evaluation, while reuse of modern bases reduces measured expenditure across the complete pipeline by 10.8%.

Read the original paper

More in Code Generation

Browse all 43 papers →
02Code Generation

Groupwise Agentic Grading and Advantage Redistribution for Code Agent RL

Jinhao Dong, Liang Zhao, Zihao Yue, Wenhan Ma, Linghao Zhang, Lei Li, Shicheng Li, Yifan Song, Bowen Ye, Fuli Luo

GAGAR helps code agents learn not only to pass tests, but to produce cleaner and more targeted implementations by redistributing RL credit according to agentic quality judgments.

Read analysis