Harness-R1: Learning to Edit Executable Runtime Harnesses from Agent Failure Trajectories
AuthorsShuai Shao, Kangning Zhang, Qingyao Li, Shijian Wang, Hao Wang, Wenxiang Jiao, Yuan Lu, Yi Guo, Weiwen Liu, Weinan Zhang
Resources
Harness-R1 trains an engineer agent to patch an LLM agent’s runtime harness from its failures, improving task success across web, household, and database environments.
Key results
Equal-weight average across WebShop, ALFWorld, and DBBench for Qwen3.5-9B.
Harness-R1 result on the frozen Qwen3.5-9B target.
Reported as 9.3 percentage points over the vanilla target.
Additional percentage-point gain from 59.2% to 64.2% after target-agent fine-tuning.
Average benchmark gain across 20 unseen target models.
Improvement on 1,270 held-out tasks, with ±1.5 percentage points across three seeds.
What the paper found
Harness-R1 treats an AI agent’s runtime harness—not just its model weights—as a trainable object. A dedicated 9B harness engineer receives batches of failure trajectories from a frozen Qwen3.5-9B agent, then generates validated executable patches across four lifecycle hooks: initialization, pre-decision guidance, pre-action mediation, and post-feedback recovery. The engineer is initialized with supervised fine-tuning from GPT-5.5-generated patches and optimized with online group-relative policy optimization, sampling 8 candidate patches per failure packet and rewarding only the target agent’s same-batch task improvement. Across WebShop, ALFWorld, and DBBench, average success rises from 44.3% to 53.6%, a gain of 9.3 percentage points, outperforming fixed or prompted editors including OpenAI’s GPT-5.5, DeepSeek-V4-Pro, Google DeepMind’s Gemini-3.5-Flash, and GLM-5.2. After direct fine-tuning of the target agent, Harness-R1 improves performance further from 59.2% to 64.2%, adding 5.0 percentage points. The learned editor transfers to 20 unseen target models with an average gain of 7.06 percentage points, and sparse evidence generalizes to 1,270 held-out tasks with an 8.9 ± 1.5 percentage-point improvement. The results position harness editing as a complementary learning axis that can co-evolve with model fine-tuning, while highlighting the importance of outcome-verified, failure-conditioned executable interventions over plausible but untested code changes.
Original abstract
Agents built around large language models continually accumulate interaction trajectories during deployment, yet their behavior typically remains fixed. Beyond updating model weights, these trajectories can improve the agent harness that constructs context, mediates tools, validates actions, and recovers execution. We introduce Harness-R1, the first method, to our knowledge, that makes failure-conditioned, lifecycle-wide editing of an existing executable runtime a learned capability. It post-trains a dedicated harness engineer with online reinforcement learning so that its edits are optimized for the realized task success they produce, rather than proposed by a fixed editor. A separate 9B engineer converts batches of target-agent failures into validated executable patches; fresh same-batch reruns of the frozen target provide outcome rewards, so training updates only the engineer. Cold-start supervised fine-tuning initializes this editing policy, which is then trained online with group-relative policy optimization. Across WebShop, ALFWorld, and DBBench, Harness-R1 raises vanilla Qwen3.5-9B success from 44.3% to 53.6% (+9.3 percentage points). After direct target-agent fine-tuning, a target-specific engineer raises the average further from 59.2% to 64.2% (+5.0 points); because these gains hold both before and after fine-tuning the target, Harness-R1 points toward co-evolving the harness engineer and the target agent.
Read the original paperMore in AI Agents
Browse all 56 papers →LEGO-Anything: Coding Agents for 3D Scene Reconstruction
Xirui Li, Peng Shi, Mingwen Dong, Sheng Zhang, Zhuoyan Xu, Dongkyu Lee, Shuaichen Chang, Yi Xiang, Lin Pan, Jiarong Jiang
LEGO-Anything turns images into editable Blender programs through iterative coding agents, offering a promising but still imperfect route to reconstructable 3D worlds.
MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution
Prithwish Jana, Mononito Goswami, Hao Liu, Xinyu Li, Langlin Huang, Zhehui Huang, Zhishen Huang, Patrick Blöbaum, Anoop Deoras, Purak Jain, Nikos Kanakaris, Sahika Genc
MILO uses teams of evolving AI agents to automatically discover better harnesses for long-horizon problem-solving systems.
Self-Organizing Agent Teams Learn to Reason Together
Aneesh Pappu, Mirac Suzgun, Yongchan Kwon, Federico Bianchi, Batu El, Mykel J. Kochenderfer, Hancheng Cao, James Zou
This work trains AI agents to discover how to divide labor, challenge ideas, and combine reasoning so that teams can solve problems no individual agent could solve alone.