DarwinX: Evolving Agent Harnesses Through Natural Selection
AuthorsYifan Zhang, Yutong Dai, Juntao Tan, Luyu Yang, Rishi Mullur, Thai Hoang, Zhiyuan Hu, James Zhu, Phil Mui, Silvio Savarese, Ran Xu, Zeyuan Chen
Resources
DarwinX improves AI agents by evolving their prompts, tools, and workflows through population-based selection rather than changing the underlying model.
Key results
Monet with DarwinX on a frozen GPT-5.5 base.
Point improvement from the 75.5% base Monet score to 83.2%.
DarwinX score using GPT-5.6 Sol.
Performance on 41 held-out tasks using Opus 4.8.
Performance on 1,260 real tasks after evolution on synthetic intents.
Official pass@1 on 500 issues using a Terminal-Bench 2.1 harness without in-domain feedback.
What the paper found
DarwinX treats agent improvement as natural selection over harnesses rather than training: the underlying model stays frozen while prompts, skills, tools, memory, and control flow evolve. Its preserve-and-extend contract accepts a variant only when it adds verified task coverage without exceeding a bounded regression, while an archive retains alternative lineages for recombination; failure-derived diagnosis, teacher trajectories, and self-contrast all feed the same harness-edit interface. On Terminal-Bench 2.1 with GPT-5.5, Monet rises from 75.5% to 83.2%, a 7.7-point gain, while GPT-5.6 Sol reaches 84.7%, ahead of Claude Code at 83.8% and Codex at 83.1%. Transfer is substantial: on 41 held-out TerminalWorld tasks using Opus 4.8, the evolved harness reaches 68.3%, exceeding the evaluated off-the-shelf agents; on WebArena-Infinity, evolution from 300 synthetic intents raises audit-clean pass@1 on 1,260 real tasks from 43.5% to 93.0%, while invalid trajectories fall from 293 to 17. A Terminal-Bench 2.1 harness transferred unchanged to 500 SWE-bench Verified issues achieves 84.2% official pass@1 without SWE-bench feedback. The strongest recurring mechanism is verification-before-finalization: explicit acceptance contracts, real-tool grounding, and persistence checks, suggesting that durable agent competence can emerge from selection over procedures even when GPT-5.5, GPT-5.6 Sol, or Opus 4.8 weights never change.
Original abstract
An LLM agent's capability depends not only on model weights but on its harness: prompts, tools, skills, and control flow. Self-improvement loops already edit harnesses, yet single-lineage search is path-dependent and local wins often regress other tasks. We introduce DarwinX, which treats self-evolution as selection over a population of harnesses with the model frozen: a preserve-and-extend contract admits only variants that extend coverage without regressing, an archive keeps alternative lineages for recombination, and failure-, teacher-, and self-derived evidence share one edit interface. Fitness comes from each benchmark's own verifier: no gold solutions, no hand-picked winners. Across four benchmarks that progressively separate the evolution signal from the test, one loop adds about 17 points on average: Terminal-Bench 2.1 rises +7.7 to 83.2% on a matched base and to the verified frontier at 84.7% on a stronger one; TerminalWorld's held-out split reaches 68.3%, ahead of every off-the-shelf agent; WebArena-Infinity real-task pass@1 rises from 43.5% to 93.0% audit-clean; and a Terminal-Bench 2.1 harness transfers unchanged to SWE-bench Verified. What evolves is general agent competence, not benchmark-specific patches, so it survives changes of task, verifier, and base model. A frozen model need not be a fixed agent: harness selection turns evaluation compute into durable capability.
Read the original paperMore in AI Agents
Browse all 56 papers →LEGO-Anything: Coding Agents for 3D Scene Reconstruction
Xirui Li, Peng Shi, Mingwen Dong, Sheng Zhang, Zhuoyan Xu, Dongkyu Lee, Shuaichen Chang, Yi Xiang, Lin Pan, Jiarong Jiang
LEGO-Anything turns images into editable Blender programs through iterative coding agents, offering a promising but still imperfect route to reconstructable 3D worlds.
MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution
Prithwish Jana, Mononito Goswami, Hao Liu, Xinyu Li, Langlin Huang, Zhehui Huang, Zhishen Huang, Patrick Blöbaum, Anoop Deoras, Purak Jain, Nikos Kanakaris, Sahika Genc
MILO uses teams of evolving AI agents to automatically discover better harnesses for long-horizon problem-solving systems.
Self-Organizing Agent Teams Learn to Reason Together
Aneesh Pappu, Mirac Suzgun, Yongchan Kwon, Federico Bianchi, Batu El, Mykel J. Kochenderfer, Hancheng Cao, James Zou
This work trains AI agents to discover how to divide labor, challenge ideas, and combine reasoning so that teams can solve problems no individual agent could solve alone.