NTH

Harness-R1: Learning to Edit Executable Runtime Harnesses from Agent Failure Trajectories

AuthorsShuai Shao, Kangning Zhang, Qingyao Li, Shijian Wang, Hao Wang, Wenxiang Jiao, Yuan Lu, Yi Guo, Weiwen Liu, Weinan Zhang

August 10, 2026 2 min read
Watch on YouTube
The one-line take

Harness-R1 trains an engineer agent to patch an LLM agent’s runtime harness from its failures, improving task success across web, household, and database environments.

Key results

44.3%
Vanilla average success before

Equal-weight average across WebShop, ALFWorld, and DBBench for Qwen3.5-9B.

53.6%
Vanilla average success after Harness-R1

Harness-R1 result on the frozen Qwen3.5-9B target.

9.3%
Vanilla improvement

Reported as 9.3 percentage points over the vanilla target.

5.0%
Fine-tuned target improvement

Additional percentage-point gain from 59.2% to 64.2% after target-agent fine-tuning.

7.06%
Unseen-target transfer gain

Average benchmark gain across 20 unseen target models.

8.9%
Held-out task gain

Improvement on 1,270 held-out tasks, with ±1.5 percentage points across three seeds.

What the paper found

Harness-R1 treats an AI agent’s runtime harness—not just its model weights—as a trainable object. A dedicated 9B harness engineer receives batches of failure trajectories from a frozen Qwen3.5-9B agent, then generates validated executable patches across four lifecycle hooks: initialization, pre-decision guidance, pre-action mediation, and post-feedback recovery. The engineer is initialized with supervised fine-tuning from GPT-5.5-generated patches and optimized with online group-relative policy optimization, sampling 8 candidate patches per failure packet and rewarding only the target agent’s same-batch task improvement. Across WebShop, ALFWorld, and DBBench, average success rises from 44.3% to 53.6%, a gain of 9.3 percentage points, outperforming fixed or prompted editors including OpenAI’s GPT-5.5, DeepSeek-V4-Pro, Google DeepMind’s Gemini-3.5-Flash, and GLM-5.2. After direct fine-tuning of the target agent, Harness-R1 improves performance further from 59.2% to 64.2%, adding 5.0 percentage points. The learned editor transfers to 20 unseen target models with an average gain of 7.06 percentage points, and sparse evidence generalizes to 1,270 held-out tasks with an 8.9 ± 1.5 percentage-point improvement. The results position harness editing as a complementary learning axis that can co-evolve with model fine-tuning, while highlighting the importance of outcome-verified, failure-conditioned executable interventions over plausible but untested code changes.

Original abstract

Agents built around large language models continually accumulate interaction trajectories during deployment, yet their behavior typically remains fixed. Beyond updating model weights, these trajectories can improve the agent harness that constructs context, mediates tools, validates actions, and recovers execution. We introduce Harness-R1, the first method, to our knowledge, that makes failure-conditioned, lifecycle-wide editing of an existing executable runtime a learned capability. It post-trains a dedicated harness engineer with online reinforcement learning so that its edits are optimized for the realized task success they produce, rather than proposed by a fixed editor. A separate 9B engineer converts batches of target-agent failures into validated executable patches; fresh same-batch reruns of the frozen target provide outcome rewards, so training updates only the engineer. Cold-start supervised fine-tuning initializes this editing policy, which is then trained online with group-relative policy optimization. Across WebShop, ALFWorld, and DBBench, Harness-R1 raises vanilla Qwen3.5-9B success from 44.3% to 53.6% (+9.3 percentage points). After direct target-agent fine-tuning, a target-specific engineer raises the average further from 59.2% to 64.2% (+5.0 points); because these gains hold both before and after fine-tuning the target, Harness-R1 points toward co-evolving the harness engineer and the target agent.

Read the original paper

More in AI Agents

Browse all 56 papers →
01Agent

LEGO-Anything: Coding Agents for 3D Scene Reconstruction

Xirui Li, Peng Shi, Mingwen Dong, Sheng Zhang, Zhuoyan Xu, Dongkyu Lee, Shuaichen Chang, Yi Xiang, Lin Pan, Jiarong Jiang

LEGO-Anything turns images into editable Blender programs through iterative coding agents, offering a promising but still imperfect route to reconstructable 3D worlds.

Read analysis
02Agent

MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution

Prithwish Jana, Mononito Goswami, Hao Liu, Xinyu Li, Langlin Huang, Zhehui Huang, Zhishen Huang, Patrick Blöbaum, Anoop Deoras, Purak Jain, Nikos Kanakaris, Sahika Genc

MILO uses teams of evolving AI agents to automatically discover better harnesses for long-horizon problem-solving systems.

Read analysis
03Agent

Self-Organizing Agent Teams Learn to Reason Together

Aneesh Pappu, Mirac Suzgun, Yongchan Kwon, Federico Bianchi, Batu El, Mykel J. Kochenderfer, Hancheng Cao, James Zou

This work trains AI agents to discover how to divide labor, challenge ideas, and combine reasoning so that teams can solve problems no individual agent could solve alone.

Read analysis