RoboCoach: World Models as Active Coaches for Compositional Robot Skills
AuthorsJiajun Liu, Yifan Chen, Yichao Liu, Jiayi Zhang, Ruoqu Chen, Shaoxuan Xie, Guocai Yao, Mengdi Xu, Sen Cui, Changshui Zhang
Affiliations[
Resources
RoboCoach uses imagined robot failures to decide what demonstrations to request next, making long-horizon manipulation skills improve more efficiently.
Key results
Spearman correlation across 22 task–policy pairs.
Additional subtask demonstrations raise Franka success to 75.0%.
Success after 150 additional subtask demonstrations.
Task-averaged success after three coaching rounds.
Average success across four unseen compositions, versus 0.0% for the shared-policy baseline.
What the paper found
RoboCoach turns a robot world model into an active teacher rather than a passive simulator. Its Route–Imagine–Diagnose–Improve, or RIDI, loop decomposes long-horizon instructions into reusable atomic skills, routes each subtask to a modular expert, rolls that expert forward in CoachWorld, and uses the RoboMeter progress judge to identify the first unresolved subtask. Recurring imagined failures determine which demonstrations to collect and which expert-specific LoRA adapter to update, while the shared VLA backbones remain frozen. CoachWorld uses conditional flow matching, calibrated two-slot end-effector trajectories, and camera-aware conditioning; it is initialized from Wan2.2 TI2V-5B. Experiments cover LIBERO, RoboTwin 2.0, and Franka and AgileX robots, using π0.5 and MolmoAct2 as policy backbones, with training conducted on eight NVIDIA H100 GPUs. Across 22 task–policy pairs, imagined and deployed success correlate at Spearman ρ = 0.840. With 150 additional subtask demonstrations, real-robot success increases from 13.3% to 75.0% on Franka and from 40.0% to 83.8% on AgileX. Final simulation success reaches 71.2% on LIBERO and 68.0% on RoboTwin 2.0. More importantly, the coached experts recombine on four held-out task compositions with 35.0% average success, compared with 0.0% for a uniformly trained shared-policy baseline, showing that imagined failure attribution can target scarce supervision at transferable skill bottlenecks.
Original abstract
Long-horizon robot manipulation reuses skills across many task compositions, but improving these compositions with additional end-to-end demonstrations is costly. A practical self-improving system must decide both what to teach next and where to apply that supervision. We present ROBOCOACH, a world-model-guided coaching framework that uses imagined failures to guide demonstration requests and expert updates. Its Route-Imagine-Diagnose-Improve (RIDI) loop executes reusable skill experts inside COACHWORLD, our shared action-conditioned world model, and uses a progress judge to record the first subtask that fails to complete. Aggregated records select which subtask demonstrations to acquire and which expert adapters to update. Across two simulation suites and two real-robot platforms, imagined and deployed success correlate over 22 task-policy pairs (rho = 0.840). Controlled comparisons show that our coaching method outperforms matched baselines under matched data budgets and update schedules. With only 150 additional subtask demonstrations, success rises from 13.3% to 75.0% on Franka and from 40.0% to 83.8% on AgileX. The coached experts also transfer to four held-out compositions, achieving an average success of 35.0%, compared with 0% for a shared-policy baseline updated with uniformly acquired demonstrations. Together, these results show that world models can serve as active coaches, turning imagined failures into targeted supervision for modular policy improvement. Project Page: https://robocoach-ai.github.io/
Read the original paperMore in World Models
Browse all 41 papers →4Director: Controlling Video World Models with Rigid 3D Geometry
Wei Cao, Hao Zhang, Vikram Voleti, Yuqun Wu, Mallikarjun B R, Shimon Vainer, Mark Boss, Yaoyao Liu
4Director makes video world models controllable by moving explicit 3D meshes through time while preserving realistic, consistent appearances.
EMPIRIC: Experiment-Driven Learning of Residual World Models for Robot Planning
Yichao Liang, Amber Li, Dat Nguyen, Emily Bunnapradist, Michelangelo Naim, Sreela Kodali, Matteo Merler, Bowen Li, Kiran Gopinathan, Yiyun Liu, Nikhil Pimpalkhare, Joshua B. Tenenbaum, Adrian Weller, Zenna Tavares, Tom Silver, Kevin Ellis
EMPIRIC lets robots discover missing physics through targeted experiments and use the resulting interpretable world models to plan better.
HappyWorld-Bench
Zhiqi Bai, Junai Cai, Yixin Chen, Jingrun Du, Tao Feng, Wei Gong, Siyuan Huang, Xiao Lin, Jiaheng Liu, Jun Luo, Yongzhe Lyu, Liya Ma, Zenan Meng, Lin Qu, Wenbo Su, Jiaming Wang, Qinghe Wang, Shaofei Wang, Yanghai Wang, Zequn Wang, Ziming Wang, Hu Wei, Jiangtao Wu, Ruiqi Wu, Jiaxin Xie, Yuchi Xu, Ze Xu, Chengting Yu, Liangyu Yuan, Gang Zeng, Yawen Zeng, Xingyao Zhang, Zizheng Zhang, Bo Zheng, Jiancheng Zhu, Song-Chun Zhu
HappyWorld-Bench tests whether AI-generated worlds remain coherent, editable, and responsive when agents explore and act within them.