NTH

Efficiency Matters in Autonomous Research

AuthorsHaiqian Yang, Yuan Cao

August 2, 2026 3 min read
Watch on YouTube
The one-line take

This paper argues that autonomous research systems should be judged not only by what they discover, but also by how efficiently they find it.

Key results

12
AutoLab task count

Systems-optimization tasks used to compare search policies.

500
Evaluation budget

Fixed candidate evaluations allocated to each policy and task.

64
Matched frontier width

Maximum number of live candidates maintained by each fixed policy and fluid.

0.780
Fluid normalized AUC

Mean efficiency score across the 12 AutoLab tasks.

0.718
Beam normalized AUC

Best-performing fixed search policy in aggregate.

0.784
Per-task oracle normalized AUC

Reference score obtained by selecting the best fixed policy separately for each task.

What the paper found

“Efficiency Matters in Autonomous Research,” by Haiqian Yang and Yuan Cao, argues that autonomous research systems should be judged not only by their final discovery but by how quickly they reach useful results under a limited evaluation budget. The paper defines search efficiency as the area under the Pareto frontier, or AUC, which rewards systems that achieve high-quality solutions early, and evaluates it alongside final reward. On 12 AutoLab systems-optimization tasks, the authors hold the verifier and Qwen3-Coder-480B proposer constant while comparing greedy hill climbing, beam search, tree search, and evolutionary search, each with frontier width 64, a budget of 500 evaluations, and three random seeds. No fixed strategy wins universally: beam search often starts fastest, while evolutionary search is more robust on crash-prone tasks. The proposed fluid method combines multi-start hill-climbing chains, an archive of diverse incumbents, crash repair, reseeding, and a UCB portfolio bandit that reallocates evaluations online according to improvement rate and incumbent quality. Fluid achieves a normalized AUC of 0.780, compared with 0.718 for the strongest fixed policy, beam search, and 0.784 for an oracle that knows the best policy per task in advance. The result extends the efficiency perspective associated with systems such as FunSearch and Google DeepMind’s AlphaEvolve: adaptive budget allocation can make autonomous research nearly oracle-efficient without knowing the task’s ideal search structure beforehand.

Original abstract

AI-driven autonomous research (AR) systems are becoming increasingly effective across a broad range of tasks. Their performance, however, is still evaluated primarily by the quality of the final outcome. In this paper, we argue that the efficiency of the solution-search process is an equally important but often overlooked dimension of performance. A strong AR system should not only produce high-quality results, but also reach them with as small a budget as possible. Search efficiency will become increasingly important as AR expands from domains with inexpensive verification, such as mathematics and coding, to real-world scientific settings in which solution evaluation may require costly physical experiments. To capture this dimension, we propose evaluating AR systems using the area under the curve (AUC) of the Pareto frontier, alongside final outcome quality. We compare several families of search algorithms, including hill climbing, beam search, tree search, and evolutionary search, across twelve systems-optimization tasks. We find that no single search structure is consistently the most efficient. We also show that search efficiency and final outcome quality are distinct performance dimensions: a method that eventually achieves the best result may nevertheless improve slowly and consume substantially more evaluation budget before reaching that result. Because the most effective search policy is generally unknown in advance, we introduce an adaptive procedure called fluid search, which uses a portfolio bandit to dynamically allocate a fixed evaluation budget across a forest of search processes. Across the evaluated tasks, fluid search achieves the highest overall search efficiency, closely matching the performance of a per-task oracle that is given the best search structure for each task in advance.

Read the original paper

More in Efficient AI

Browse all 55 papers →
01Efficiency

Decoding Looped Transformers Better for (Almost) Free

Weihao Liu, Huangjie Zheng, Tianrong Chen, Rohit Dilip, Richard He Bai, Yizhu Jiao, Yuyang Wang, Ruixiang Zhang

LoopCD turns the partially computed states of looped Transformers into free guidance, improving accuracy while often cutting inference compute nearly in half.

Read analysis
02Efficiency

Scaling Laws for Looped Mixture of Experts

Yanbei Chen, Anirudh Goyal, Raghuraman Krishnamoorthi

This work develops scaling laws that explain how looping and sparse experts can be combined to build more capable models with less training and inference compute.

Read analysis