Scaling Automatic Research Agents via World Models
AuthorsXiyuan Yang, Sheikh Sarwar, Jingru Cheng, Zhan Shi, Duanshun Li, Huiyuan Chen, Haiyang Zhang, Xing Fan, Chenlei Guo, Jingrui He, Zhenyu Liao
Resources
This work speeds up the training of autonomous research agents by letting them practice in a learned simulated world instead of repeatedly running costly real experiments.
Key results
WMRL reduces real-execution GRPO compute by 3.1 times for Qwen3.5-4B.
WMRL reduces real-execution GRPO compute by 3.4 times for Qwen3.5-9B.
Average held-out MLE-Dojo leaderboard percentile for the 4B WMRL agent.
Average held-out MLE-Dojo leaderboard percentile for the 9B WMRL agent.
Overall success rate for MiniVLA-1B trained with WMRL.
What the paper found
This Amazon-backed paper introduces World Model Reinforcement Learning, or WMRL, for automatic research agents whose reinforcement-learning cost is dominated by running each generated solution in an isolated GPU sandbox rather than by language-model generation. WMRL replaces most real execution with a prompted, same-backbone language-model world model, then uses roughly 10% real-execution anchor groups to correct its errors: Online Debiasing fits a monotone isotonic-regression map that removes systematic reward bias, while Inverse-Variance Denoising fuses predicted and real rewards with variance-optimal weights. The analysis shows that uncorrected world-model rewards add O(B²) bias and O(σ²) noise terms to the convergence bound, whereas WMRL makes the bias contract over training and reduces variance below either reward stream alone. On MLE-Dojo, an interactive successor to OpenAI’s MLE-Bench, Qwen3.5-4B and Qwen3.5-9B WMRL training cuts compute by 3.1 times and 3.4 times versus real-execution GRPO, reaching 16.4% and 21.6% average held-out leaderboard percentile, respectively; the 4B and 9B agents also outperform Kimi-48B-A3B and Nemotron-120B-A12B. The approach transfers to embodied learning with MiniVLA-1B, Robometer, and LIBERO-Long, reaching 41.2% overall success, while experiments use NVIDIA A100 GPUs. Together, the results show that scalable learned execution signals can replace most expensive rollouts without sacrificing downstream performance.
Original abstract
Automating empirical research is a long-standing direction of AI. Recent automatic research (AutoResearch) agents bring this goal within reach, as modern LLMs show the capability to independently implement solutions and learn from the execution outcomes. Behind these gains, post-training (especially RL) plays a central role. In this paper, we identify a fundamental tension when scaling RL for these agents: the two components of every AutoResearch trajectory (agent generation and environment execution) scale in very different manners, since all generation shares compute through batching, while each execution occupies its exclusive sandbox and real machine time. As a result, the environment execution dominates the training cost and becomes the bottleneck as trajectories grow. To resolve this tension, we propose World Model RL (WMRL), which replaces environment execution with a world model to remove this bottleneck. Additionally, the world model can be imperfect, as its rewards are corrupted by bias and noise. Therefore, we further equip WMRL with two mitigations, Online Debiasing and Inverse-Variance Denoising, which offset the bias and suppress the noise respectively. Theoretically, we prove that both mitigations of WMRL strictly improve the convergence guarantee. Empirically, WMRL accelerates training by 3-4x on various tasks at different agent scales, while exceeding the performance of standard RL baselines. Moreover, our post-trained 4B and 9B agents outperform much larger open-weight agents of 48B and 120B on held-out benchmarks. Beyond AutoResearch, WMRL also transfers to post-training embodied VLA policies, which demonstrates the generalizability of our method.
Read the original paperMore in World Models
Browse all 41 papers →4Director: Controlling Video World Models with Rigid 3D Geometry
Wei Cao, Hao Zhang, Vikram Voleti, Yuqun Wu, Mallikarjun B R, Shimon Vainer, Mark Boss, Yaoyao Liu
4Director makes video world models controllable by moving explicit 3D meshes through time while preserving realistic, consistent appearances.
EMPIRIC: Experiment-Driven Learning of Residual World Models for Robot Planning
Yichao Liang, Amber Li, Dat Nguyen, Emily Bunnapradist, Michelangelo Naim, Sreela Kodali, Matteo Merler, Bowen Li, Kiran Gopinathan, Yiyun Liu, Nikhil Pimpalkhare, Joshua B. Tenenbaum, Adrian Weller, Zenna Tavares, Tom Silver, Kevin Ellis
EMPIRIC lets robots discover missing physics through targeted experiments and use the resulting interpretable world models to plan better.
RoboCoach: World Models as Active Coaches for Compositional Robot Skills
Jiajun Liu, Yifan Chen, Yichao Liu, Jiayi Zhang, Ruoqu Chen, Shaoxuan Xie, Guocai Yao, Mengdi Xu, Sen Cui, Changshui Zhang
RoboCoach uses imagined robot failures to decide what demonstrations to request next, making long-horizon manipulation skills improve more efficiently.