EMPIRIC: Experiment-Driven Learning of Residual World Models for Robot Planning
AuthorsYichao Liang, Amber Li, Dat Nguyen, Emily Bunnapradist, Michelangelo Naim, Sreela Kodali, Matteo Merler, Bowen Li, Kiran Gopinathan, Yiyun Liu, Nikhil Pimpalkhare, Joshua B. Tenenbaum, Adrian Weller, Zenna Tavares, Tom Silver, Kevin Ellis
AffiliationsBasis Research Institute · University of Cambridge · Carnegie Mellon University · Fondazione Bruno Kessler · Massachusetts Institute of Technology · The Alan Turing Institute
Resources
EMPIRIC lets robots discover missing physics through targeted experiments and use the resulting interpretable world models to plan better.
Key results
EMPIRIC was evaluated across five PyBullet manipulation domains.
EMPIRIC solved 100% of runs, or 25/25.
Removing parameter-belief fitting reduced performance to 20/25 runs.
Removing explicit uncertainty reduced performance to 21/25 runs.
The learned model predicted a 10.1 ± 2.3 cm domino slide.
What the paper found
EMPIRIC, or Experiment-driven Modeling of Physics: Inferring Residuals In Code, enables a robot to extend a generic rigid-body simulator with executable Python mechanisms for physics it does not know, such as glue curing, balloon lift, water heating, and airflow. A coding-agent harness—using Anthropic’s Claude Opus 5, with interfaces related to the Anthropic Claude Agent SDK and OpenAI Codex—writes and revises these mechanisms, adds hidden recurrent state, fits parameters through Bayesian-style replay inference, and plans across joint samples of uncertain parameters and noisy object states. Information-seeking experiments are selected through mutual information over learned predicates, while execution monitors predicted outcomes and triggers replanning or model revision after failures. In five PyBullet manipulation domains, EMPIRIC solved 100% of runs, or 25/25, versus 16 for the Direct agent, 14 for Direct plus scene, and 16 for Standalone simulation, while using fewer environment steps than the baselines in most comparisons. Removing parameter-belief fitting reduced performance to 20/25 runs, and removing explicit uncertainty reduced it to 21/25, especially harming irreversible balloon-release tasks. On a physical Franka Emika Panda robot, perception used SAM 2 and stereo depth; after two gust experiments, EMPIRIC learned wind decay and domino masses, predicting a 10.1 ± 2.3 cm slide against a measured 9.9 cm and solving a two-domino cascade in a new target layout.
Original abstract
A robot should be able to learn through experiments how unfamiliar objects behave and interact, then plan with that knowledge. It need not start from scratch: physics engines supply knowledge of motion and contact, but can omit entire mechanisms, such as glue curing, water heating, or wind. We present EMPIRIC, an agent that learns a residual world model: a physics engine extended with code for the missing mechanisms. The learned programs can introduce new forces, constraints, and hidden state, and Bayesian inference estimates their parameters and states from noisy observations. The resulting model lets the agent predict the outcomes of actions, choose informative experiments, and revise its hypotheses when predictions fail. Across five simulated domains, EMPIRIC learns interpretable, reusable models, and solves more tasks with fewer environment interactions than all three baselines. On a physical robot, it learns wind forces and domino masses to solve a manipulation task. Website and code: https://yichao-liang.github.io/empiric
Read the original paperMore in World Models
Browse all 41 papers →4Director: Controlling Video World Models with Rigid 3D Geometry
Wei Cao, Hao Zhang, Vikram Voleti, Yuqun Wu, Mallikarjun B R, Shimon Vainer, Mark Boss, Yaoyao Liu
4Director makes video world models controllable by moving explicit 3D meshes through time while preserving realistic, consistent appearances.
RoboCoach: World Models as Active Coaches for Compositional Robot Skills
Jiajun Liu, Yifan Chen, Yichao Liu, Jiayi Zhang, Ruoqu Chen, Shaoxuan Xie, Guocai Yao, Mengdi Xu, Sen Cui, Changshui Zhang
RoboCoach uses imagined robot failures to decide what demonstrations to request next, making long-horizon manipulation skills improve more efficiently.
HappyWorld-Bench
Zhiqi Bai, Junai Cai, Yixin Chen, Jingrun Du, Tao Feng, Wei Gong, Siyuan Huang, Xiao Lin, Jiaheng Liu, Jun Luo, Yongzhe Lyu, Liya Ma, Zenan Meng, Lin Qu, Wenbo Su, Jiaming Wang, Qinghe Wang, Shaofei Wang, Yanghai Wang, Zequn Wang, Ziming Wang, Hu Wei, Jiangtao Wu, Ruiqi Wu, Jiaxin Xie, Yuchi Xu, Ze Xu, Chengting Yu, Liangyu Yuan, Gang Zeng, Yawen Zeng, Xingyao Zhang, Zizheng Zhang, Bo Zheng, Jiancheng Zhu, Song-Chun Zhu
HappyWorld-Bench tests whether AI-generated worlds remain coherent, editable, and responsive when agents explore and act within them.