NTH

MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution

AuthorsPrithwish Jana, Mononito Goswami, Hao Liu, Xinyu Li, Langlin Huang, Zhehui Huang, Zhishen Huang, Patrick Blöbaum, Anoop Deoras, Purak Jain, Nikos Kanakaris, Sahika Genc

AffiliationsGeorgia Institute of Technology · AWS AI Labs · Carnegie Mellon University · Washington University in St. Louis

October 3, 2026 2 min read
Watch on YouTube
The one-line take

MILO uses teams of evolving AI agents to automatically discover better harnesses for long-horizon problem-solving systems.

Key results

12.0%
Terminal-Bench improvement

Resolution-rate improvement over the initial harness with Opus 4.8 on Terminal-Bench 2.1.

28.3%
PaperBench improvement

Resolution-rate improvement over the initial harness with Opus 4.8 on PaperBench.

10.3%
DeepSWE improvement

Resolution-rate improvement over the initial harness with Opus 4.8 on DeepSWE.

86.1
Terminal-Bench MILO score

Resolution rate in percent on Terminal-Bench 2.1 with Opus 4.8.

26%
Token reduction

Fewer tokens than the initial harness on Terminal-Bench 2.1 with Opus 4.8.

15.6%
gpt-oss-120b DeepSWE score

MILO resolution rate on DeepSWE with the gpt-oss-120b backbone.

What the paper found

MILO, or Meta-evolutionary Island Orchestration, automates the design of AI-agent harnesses—the control layer governing prompts, tools, memory, sub-agents, verification, and stopping—without retraining the underlying model. Its novelty is a coupled evolutionary process: island-specific lineage trees retain both successful and rejected harness mutations, evidence-driven coding agents rewrite complete harnesses from failure traces, and a meta-evolutionary orchestrator intervenes when progress stalls by reassigning mutators, grafting capabilities across islands, preserving niches through speciation, or reshaping the task curriculum. Candidates are admitted on a Pareto frontier balancing resolution rate, token use, and latency. Across Terminal-Bench 2.1, PaperBench, and DeepSWE, MILO outperformed eight engineered harnesses and six search baselines with both Opus 4.8 and gpt-oss-120b. With Opus 4.8, resolution rate improved over the initial harness by 12.0% on Terminal-Bench 2.1, 28.3% on PaperBench, and 10.3% on DeepSWE. On Terminal-Bench 2.1 it reached 86.1±2.0%, exceeding the official leaderboard’s 83.8±2.3%, while using 26% fewer tokens than the initial harness. With gpt-oss-120b, it achieved 15.6% on DeepSWE, roughly 20× its initial harness. The approach complements production-style systems such as Anthropic’s Claude Code and OpenAI’s Codex by evolving their surrounding control architecture rather than their weights, and it also improved three open mathematical bounds on EinsteinArena.

Original abstract

Modern agentic systems combine an AI model with a harness that controls execution and environmental interactions. Harness design strongly affects long-horizon performance, yet its combinatorial search space demands substantial human effort that must be repeated as models change. Existing automated methods explore this space narrowly, optimizing only components such as prompts or skills or becoming trapped by fixed, exploitative search strategies. We introduce MILO (Meta-evolutionary Island Orchestration), a framework that co-evolves agent harnesses and the strategy used to discover them. MILO combines: (i) hierarchical lineage memory over island-based trees, using rejected mutations as negative evidence; (ii) per-island mutator agents that rewrite complete harnesses using global search history and parent-specific feedback; and (iii) an orchestrator that adapts search through lineage grafting and speciation, mutator reassignment and curriculum revision. Across Terminal-Bench 2.1, PaperBench, and DeepSWE, MILO-discovered harnesses outperform eight state-of-the-art harnesses and six search methods using frontier (Opus 4.8) and open-weight (gpt-oss-120b) models. With Opus 4.8, MILO improves resolution over its initial harness by $+12.0\%$, $+28.3\%$, and $+10.3\%$, respectively, compared with best prior-search gains of $+4.5\%$, $+18.3\%$, and $0\%$. On Terminal-Bench 2.1, it achieves $86.1 \pm 2.0\%$, exceeding the official leaderboard's top entry ($83.8 \pm 2.3\%$) while using 26\% fewer tokens than its initial harness. On EinsteinArena open problems, MILO improves best-known upper bounds for Erdős minimum-overlap ($0.3808586 \to 0.3808568$) and the first and third autocorrelation inequalities ($1.50274365 \to 1.50274360$; $1.45081 \to 1.44889$).

Read the original paper

More in AI Agents

Browse all 56 papers →
01Agent

LEGO-Anything: Coding Agents for 3D Scene Reconstruction

Xirui Li, Peng Shi, Mingwen Dong, Sheng Zhang, Zhuoyan Xu, Dongkyu Lee, Shuaichen Chang, Yi Xiang, Lin Pan, Jiarong Jiang

LEGO-Anything turns images into editable Blender programs through iterative coding agents, offering a promising but still imperfect route to reconstructable 3D worlds.

Read analysis
02Agent

Self-Organizing Agent Teams Learn to Reason Together

Aneesh Pappu, Mirac Suzgun, Yongchan Kwon, Federico Bianchi, Batu El, Mykel J. Kochenderfer, Hancheng Cao, James Zou

This work trains AI agents to discover how to divide labor, challenge ideas, and combine reasoning so that teams can solve problems no individual agent could solve alone.

Read analysis
03Agent

Atria Dawn: The Dawn of Agentic Superintelligence

Honglin Guo, Tao Gui, Yicheng Chen, Guanting Dong, Qiming Ge, Yuyang Hu, Zixian Huang, Jiajie Jin, Alexander Lam, Yining Li, Jiahang Lin, Yanjiang Liu, Xinyu Lu, Haijun Lv, Junlin Shang, Qisheng Su, Guoqiang Wang, Rui Wang, Zhecan Wang, Hao Xiang, Xinchen Xie, Shuhao Xing, Xiaoyu Xing, Wanghan Xu, Xinyu Yang, Yajie Yang, Chengfeng Zhao, Haoran Zhao, Ruojun Zhou, Yunhua Zhou, Yicheng Zou, Kun Cai, Qiye Cai, Xinmeng Che, Haodong Chen, Jiabei Chen, Jiahao Chen, Jiayi Chen, Yujia Chen, Lizhi Cui, Youheng Dai, Xin Deng, Yi Dong, Shihan Dou, Chenya Gu, Xu Guo, Ding Han, Feiyang Hao, Haotan He, Jie Hou, Binze Hu, Zijian Hu, Junhao Huang, Huicheng Jiang, Jiazhen Jiang, Shufan Jiang, Jiahao Kuang, Bowen Lai, Bo Li, Jiaqiang Li, Peng Li, Qilong Li, Zhuoqun Li, Jiaxiang Liu, Shuainan Liu, Tong Liu, Yi Liu, Zhonghang Lu, Jianwen Luo, Yanyi Luo, Huijie Lv, Ningsheng Ma, Zerun Ma, Houcheng Min, Chengjun Pan, Qiyuan Peng, Xiaoxuan Peng, Jianmin Qian, Jiantao Qiu, Wanying Ren, Huayu Sha, Jifei Shan...

Atria Dawn explores how research-oriented AI agents can move beyond completing tasks to partnering with humans on scientific discovery while preserving human oversight.

Read analysis