NTH

RoboCoach: World Models as Active Coaches for Compositional Robot Skills

AuthorsJiajun Liu, Yifan Chen, Yichao Liu, Jiayi Zhang, Ruoqu Chen, Shaoxuan Xie, Guocai Yao, Mengdi Xu, Sen Cui, Changshui Zhang

Affiliations[

October 5, 2026 2 min read
Watch on YouTube
The one-line take

RoboCoach uses imagined robot failures to decide what demonstrations to request next, making long-horizon manipulation skills improve more efficiently.

Key results

0.840
Imagined-deployed success correlation

Spearman correlation across 22 task–policy pairs.

150
Franka demonstration budget

Additional subtask demonstrations raise Franka success to 75.0%.

83.8%
AgileX coached success

Success after 150 additional subtask demonstrations.

71.2%
LIBERO final success

Task-averaged success after three coaching rounds.

35.0%
Held-out composition success

Average success across four unseen compositions, versus 0.0% for the shared-policy baseline.

What the paper found

RoboCoach turns a robot world model into an active teacher rather than a passive simulator. Its Route–Imagine–Diagnose–Improve, or RIDI, loop decomposes long-horizon instructions into reusable atomic skills, routes each subtask to a modular expert, rolls that expert forward in CoachWorld, and uses the RoboMeter progress judge to identify the first unresolved subtask. Recurring imagined failures determine which demonstrations to collect and which expert-specific LoRA adapter to update, while the shared VLA backbones remain frozen. CoachWorld uses conditional flow matching, calibrated two-slot end-effector trajectories, and camera-aware conditioning; it is initialized from Wan2.2 TI2V-5B. Experiments cover LIBERO, RoboTwin 2.0, and Franka and AgileX robots, using π0.5 and MolmoAct2 as policy backbones, with training conducted on eight NVIDIA H100 GPUs. Across 22 task–policy pairs, imagined and deployed success correlate at Spearman ρ = 0.840. With 150 additional subtask demonstrations, real-robot success increases from 13.3% to 75.0% on Franka and from 40.0% to 83.8% on AgileX. Final simulation success reaches 71.2% on LIBERO and 68.0% on RoboTwin 2.0. More importantly, the coached experts recombine on four held-out task compositions with 35.0% average success, compared with 0.0% for a uniformly trained shared-policy baseline, showing that imagined failure attribution can target scarce supervision at transferable skill bottlenecks.

Original abstract

Long-horizon robot manipulation reuses skills across many task compositions, but improving these compositions with additional end-to-end demonstrations is costly. A practical self-improving system must decide both what to teach next and where to apply that supervision. We present ROBOCOACH, a world-model-guided coaching framework that uses imagined failures to guide demonstration requests and expert updates. Its Route-Imagine-Diagnose-Improve (RIDI) loop executes reusable skill experts inside COACHWORLD, our shared action-conditioned world model, and uses a progress judge to record the first subtask that fails to complete. Aggregated records select which subtask demonstrations to acquire and which expert adapters to update. Across two simulation suites and two real-robot platforms, imagined and deployed success correlate over 22 task-policy pairs (rho = 0.840). Controlled comparisons show that our coaching method outperforms matched baselines under matched data budgets and update schedules. With only 150 additional subtask demonstrations, success rises from 13.3% to 75.0% on Franka and from 40.0% to 83.8% on AgileX. The coached experts also transfer to four held-out compositions, achieving an average success of 35.0%, compared with 0% for a shared-policy baseline updated with uniformly acquired demonstrations. Together, these results show that world models can serve as active coaches, turning imagined failures into targeted supervision for modular policy improvement. Project Page: https://robocoach-ai.github.io/

Read the original paper

More in World Models

Browse all 41 papers →
01World Model

4Director: Controlling Video World Models with Rigid 3D Geometry

Wei Cao, Hao Zhang, Vikram Voleti, Yuqun Wu, Mallikarjun B R, Shimon Vainer, Mark Boss, Yaoyao Liu

4Director makes video world models controllable by moving explicit 3D meshes through time while preserving realistic, consistent appearances.

Read analysis
02World Model

EMPIRIC: Experiment-Driven Learning of Residual World Models for Robot Planning

Yichao Liang, Amber Li, Dat Nguyen, Emily Bunnapradist, Michelangelo Naim, Sreela Kodali, Matteo Merler, Bowen Li, Kiran Gopinathan, Yiyun Liu, Nikhil Pimpalkhare, Joshua B. Tenenbaum, Adrian Weller, Zenna Tavares, Tom Silver, Kevin Ellis

EMPIRIC lets robots discover missing physics through targeted experiments and use the resulting interpretable world models to plan better.

Read analysis
03World Model

HappyWorld-Bench

Zhiqi Bai, Junai Cai, Yixin Chen, Jingrun Du, Tao Feng, Wei Gong, Siyuan Huang, Xiao Lin, Jiaheng Liu, Jun Luo, Yongzhe Lyu, Liya Ma, Zenan Meng, Lin Qu, Wenbo Su, Jiaming Wang, Qinghe Wang, Shaofei Wang, Yanghai Wang, Zequn Wang, Ziming Wang, Hu Wei, Jiangtao Wu, Ruiqi Wu, Jiaxin Xie, Yuchi Xu, Ze Xu, Chengting Yu, Liangyu Yuan, Gang Zeng, Yawen Zeng, Xingyao Zhang, Zizheng Zhang, Bo Zheng, Jiancheng Zhu, Song-Chun Zhu

HappyWorld-Bench tests whether AI-generated worlds remain coherent, editable, and responsive when agents explore and act within them.

Read analysis