NTH

HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness?

AuthorsYuhao Wu, Jingyuan Zhang, Jiajun Shi, Xinping Lei, Qingshui Gu, Yuxuan Zhang, Zexuan Wang, Chen He, Chen Huang, Maojia Song, Zhiyuan Zeng, Shaowen Wang, Jinkai Liu, Yunfeng Shi, Jiaheng Liu, Shen Yan, Wenhao Huang, Ge Zhang, Wenxuan Zhang

September 4, 2026 3 min read
Watch on YouTube
The one-line take

HarnessDev asks whether LLMs can build and improve the software infrastructure that makes their own agents effective.

Key results

2,207
Creation evaluation coverage

Unique downstream instances spanning four domains and five benchmarks.

67.8
Opus 4.8 overall Self-Eval

Best overall generated-harness score, compared with the 86.2 human-engineered reference.

32.9
Opus 4.8 MLE-bench medal rate

Highest reported machine-learning experimentation medal rate, narrowly above Gemini 3.1 Pro at 32.4.

4.44
Best held-out Evolution improvement

Largest self-runtime gain on the hidden 630-task SWE-Pro split.

53.1%
Feedback-to-held-out agreement

Fraction of comparable evolution switches where feedback and held-out scores moved in the same direction.

What the paper found

HarnessDev reframes agent evaluation: instead of scoring only answers produced inside a fixed scaffold, it evaluates whether large language models can build and evolve the scaffold itself—the agent harness governing execution loops, tools, context, state, recovery, and verification. In Creation, six models, including Anthropic’s Claude Opus 4.8, OpenAI’s GPT-5.5, Google’s Gemini 3.1 Pro, DeepSeek V4 Pro, Qwen 3.7 Max, and ByteDance’s Seed 2.0 Pro, construct runnable harnesses from a deliberately weak seed across four domains and five benchmarks totaling 2,207 downstream instances. Opus 4.8 achieved the strongest overall Self-Eval score at 67.8, still below the selected human-engineered reference at 86.2; generated systems performed relatively well on writing and MLE-bench, where Opus reached a 32.9 medal rate, but lagged substantially on code, search, and research. The study also shows that execution cost and capability are weakly coupled, and that harnesses often co-adapt to their creator’s runtime model: Opus’s SWE-Pro score fell from 69.3 to 33.0 when executed by Gemini 3.1 Pro. In Evolution, models revised their own code harnesses using feedback from 100 SWE-bench Pro tasks and 89 Terminal-Bench 2.1 tasks, then faced a hidden 630-task SWE-Pro split. Self-runtime evolution produced modest held-out gains, with the best improvement reaching 4.44 points, but feedback and held-out scores moved together only 53.1% of the time. The conclusion is that LLMs can make useful local harness repairs, but robust, transferable, regression-safe self-development remains unsolved.

Original abstract

As agents move from research prototypes to deployed tools, their capability increasingly depends on model-external execution infrastructure, commonly termed the agent harness. Changing this harness while holding model weights fixed can substantially alter task performance. Current agent evaluations typically report downstream performance under a chosen harness, leaving a model's ability to develop the harness itself comparatively underexplored. We introduce HarnessDev, a benchmark that shifts the unit of evaluation from task outputs to runnable infrastructure. HarnessDev covers two stages. In Creation, the agent starts from a minimal seed and a small number of cases, then builds a complete execution system. In Evolution, it starts from its own created harness and iteratively revises it using downstream execution feedback, with the goal of improving benchmark performance. We then evaluate each constructed harness on capability (task success on held-out benchmarks) and efficiency (execution-token cost). The reported Creation results cover six creator LLMs, four domains, and five downstream benchmarks totaling 2,207 unique downstream instances, with hidden evaluation tasks withheld from development. We find that generated harnesses remain substantially behind mature human-engineered references on code and on search and research, while matching or exceeding the selected references on writing and machine-learning experimentation, with large variation in execution cost. Evolution produces some performance gains, but they are unstable and transfer only partially to held-out tasks. Experiments with a fixed runtime model further show that the gains depend strongly on the model executing the harness, indicating limited transfer across models.

Read the original paper

More in AI Agents

Browse all 56 papers →
01Agent

LEGO-Anything: Coding Agents for 3D Scene Reconstruction

Xirui Li, Peng Shi, Mingwen Dong, Sheng Zhang, Zhuoyan Xu, Dongkyu Lee, Shuaichen Chang, Yi Xiang, Lin Pan, Jiarong Jiang

LEGO-Anything turns images into editable Blender programs through iterative coding agents, offering a promising but still imperfect route to reconstructable 3D worlds.

Read analysis
02Agent

MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution

Prithwish Jana, Mononito Goswami, Hao Liu, Xinyu Li, Langlin Huang, Zhehui Huang, Zhishen Huang, Patrick Blöbaum, Anoop Deoras, Purak Jain, Nikos Kanakaris, Sahika Genc

MILO uses teams of evolving AI agents to automatically discover better harnesses for long-horizon problem-solving systems.

Read analysis
03Agent

Self-Organizing Agent Teams Learn to Reason Together

Aneesh Pappu, Mirac Suzgun, Yongchan Kwon, Federico Bianchi, Batu El, Mykel J. Kochenderfer, Hancheng Cao, James Zou

This work trains AI agents to discover how to divide labor, challenge ideas, and combine reasoning so that teams can solve problems no individual agent could solve alone.

Read analysis