NTH

Toward Generalist Autonomous Research via Hypothesis-Tree Refinement

AuthorsJiajie Jin, Yuyang Hu, Kai Qiu, Qi Dai, Chong Luo, Guanting Dong, Xiaoxi Li, Tong Zhao, Xiaolong Ma, Gongrui Zhang, Zhirong Wu, Bei Liu, Zhengyuan Yang, Linjie Li, Lijuan Wang, Hongjin Qian, Yutao Zhu, Zhicheng Dou

July 13, 2026 3 min read
Watch on YouTube
The one-line take

This paper introduces Arbor, an AI research agent that keeps a living tree of hypotheses and evidence so it can iteratively plan, test, and improve scientific ideas over long time horizons.

Key results

2.5x
average relative held-out gain vs Codex/Claude Code

Arbor exceeds the average held-out gain of both baselines under the same budget

77.36
Terminal-Bench 2.0 held-out pass rate

Best held-out result on the harness-engineering task

67.67
BrowseComp held-out accuracy

Best held-out result on the browsing benchmark

18.00
Search-Agent Data Synthesis held-out gap

Best held-out result on the search-agent synthesis task

20.83
Math-Reasoning Data Synthesis held-out gap

Best held-out result on the math synthesis task

86.36
MLE-Bench Lite Any Medal

Arbor with GPT-5.5 achieves the strongest reported MLE-Bench Lite result

What the paper found

Toward Generalist Autonomous Research via Hypothesis-Tree Refinement introduces Arbor, a long-horizon autonomous research system built at Renmin University of China and Microsoft Research that formalizes autonomous optimization as iterative artifact improvement under dev/test separation. Arbor’s core mechanism is Hypothesis Tree Refinement, a persistent tree that stores hypotheses, executable artifact branches, experimental evidence, and distilled insights, while a long-lived coordinator manages global strategy and short-lived executors test individual hypotheses in isolated git worktrees. Across six real research tasks spanning model training, harness engineering, and data synthesis, Arbor reaches the best held-out result on all six and delivers more than 2.5× the average relative held-out gain of Codex and Claude Code under the same interface and resource budget. The strongest single-task numbers include 77.36% on Terminal-Bench 2.0, 67.67% on BrowseComp, 18.00 on Search-Agent Data Synthesis, and 20.83 on Math-Reasoning Data Synthesis, while on MLE-Bench Lite Arbor with GPT-5.5 attains 86.36% Any Medal, the best result in the comparison. Ablations show the tree is not just bookkeeping: removing the tree or disabling insight propagation drops Any Medal from 81.82% to 63.64% and 54.54% with the Claude Opus 4.6 backbone. The paper’s technical novelty is that verified progress is admitted only through a held-out merge gate, so dev-set gains become evidence rather than final success unless they transfer to the test evaluator.

Original abstract

Scientific progress depends on a repeated loop of exploration, experimentation, and abstraction. Researchers test candidate directions, interpret the evidence, and carry the resulting lessons into later attempts. We study how an AI agent can run this loop autonomously over long horizons. We introduce Arbor, a general framework for autonomous research that combines a long-lived coordinator, short-lived executors, and Hypothesis Tree Refinement (HTR), a persistent tree that links hypotheses, artifacts, evidence, and distilled insights across time. The coordinator manages global research strategy over the tree, while executors implement and test individual hypotheses in isolated worktrees. As results return, Arbor updates the tree, propagates reusable lessons, refines the search frontier, and admits verified improvements. This design turns autonomous research from a sequence of local attempts into a cumulative process in which strategy, execution, and evidence are carried across time. We evaluate Arbor under Autonomous Optimization (AO), an operational setting where an agent improves an initial research artifact through iterative experimentation without step-level human supervision. Across six real research tasks in model training, harness engineering, and data synthesis, Arbor achieves the best held-out result on all six tasks, attaining more than 2.5x the average relative held-out gain of Codex and Claude Code under the same task interface and resource budget. On MLE-Bench Lite, Arbor reaches 86.36% Any Medal with GPT-5.5, the strongest result in our comparison.

Read the original paper

More in AI Agents

Browse all 56 papers →
01Agent

LEGO-Anything: Coding Agents for 3D Scene Reconstruction

Xirui Li, Peng Shi, Mingwen Dong, Sheng Zhang, Zhuoyan Xu, Dongkyu Lee, Shuaichen Chang, Yi Xiang, Lin Pan, Jiarong Jiang

LEGO-Anything turns images into editable Blender programs through iterative coding agents, offering a promising but still imperfect route to reconstructable 3D worlds.

Read analysis
02Agent

MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution

Prithwish Jana, Mononito Goswami, Hao Liu, Xinyu Li, Langlin Huang, Zhehui Huang, Zhishen Huang, Patrick Blöbaum, Anoop Deoras, Purak Jain, Nikos Kanakaris, Sahika Genc

MILO uses teams of evolving AI agents to automatically discover better harnesses for long-horizon problem-solving systems.

Read analysis
03Agent

Self-Organizing Agent Teams Learn to Reason Together

Aneesh Pappu, Mirac Suzgun, Yongchan Kwon, Federico Bianchi, Batu El, Mykel J. Kochenderfer, Hancheng Cao, James Zou

This work trains AI agents to discover how to divide labor, challenge ideas, and combine reasoning so that teams can solve problems no individual agent could solve alone.

Read analysis