WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning
AuthorsSenyu Fei, Xiaopeng Yu, Siyin Wang, Xianzhong Zhao, Jingjing Gong, Xipeng Qiu
Resources
WCM improves robot-learning critics by predicting how the world evolves over time instead of relying only on single-frame value estimates.
Key results
Tasks evaluated across four manipulation benchmarks.
WCM result under in-distribution evaluation.
WCM result under out-of-distribution evaluation.
Improvement from 0.8% zero-shot success to 98.7% after WCM training.
Learnable parameters used in the seven-task WidowX-250S experiments.
What the paper found
WCM, from researchers at Tongji University, Shanghai Innovation Institute, and Fudan University, addresses a central weakness in vision-language-action reinforcement learning: conventional critics estimate value from a single image or weakly supervised frame history, even though robot control is a partially observable process requiring motion and contact information. The World Critic Model uses a lightweight LeJEPA architecture with a causal Transformer history trunk, CLIP language conditioning, and action-conditioned gated FiLM blocks to jointly predict the next latent state and estimate return. Its objective combines value regression, future-latent prediction, and SIGReg regularization, producing a predictive state representation rather than merely fitting scalar rewards. WCM integrates with PPO and Flow-SDE for on-policy training and AWR or RECAP for off-policy training, and supports Physical Intelligence’s π0 and π0.5 models as well as OpenVLA-OFT. Across 149 tasks and four simulation benchmarks, it improves both in-distribution and out-of-distribution performance; with OpenVLA-OFT on ManiSkill, WCM reaches 99.0% in-distribution success and 77.9% out-of-distribution success. Starting from a 0.8% zero-shot baseline, it achieves 98.7%, a 12,551% improvement. On seven WidowX-250S real-world manipulation tasks, WCM uses 107.2M learnable parameters and consistently outperforms standard critics, including on cloth and towel folding, stovetop cleaning, and rotating-sushi picking. Ablations show that world prediction—not history alone—is the key contributor, while a history length of three frames performs best on average.
Original abstract
Reinforcement learning (RL) post-training of Vision-Language-Action (VLA) models has shown strong promise for robotic manipulation. Among RL methods, critic-based approaches rely on a value estimator that predominantly operates on single-frame observations or single-frame VLM backbone latents, which is a fundamental mismatch with the partially observable nature of robot control. A naive approach to incorporate observation history into the critic incurs exponential complexity with high-dimensional visual space, and still fails because pure scalar-return regression provides insufficient supervision for learning cross-temporal dynamics. We identify the root cause as a state approximation problem: without an explicit world modeling objective, the critic's representation cannot capture the temporal structure needed for accurate value estimation. To address this, we propose the World Critic Model (WCM), built on a lightweight LeJEPA architecture; WCM jointly predicts future latent state and estimates values, such that the critic's representation is explicitly trained to capture temporal dynamics rather than merely regress scalar returns. WCM integrates seamlessly into both on-policy and off-policy training pipelines and is compatible with state-of-the-art VLA backbones including Pi0, Pi0.5, and OpenVLA-OFT. Extensive experiments on 149 tasks across four benchmarks demonstrate that WCM consistently achieves state-of-the-art performance in both in-distribution and out-of-distribution settings, with particularly strong generalization gains. We further validate WCM on seven real-world manipulation tasks using OpenVLA-OFT and Pi0.5 with off-policy RL, confirming stable deployment across diverse settings.
Read the original paperMore in World Models
Browse all 41 papers →4Director: Controlling Video World Models with Rigid 3D Geometry
Wei Cao, Hao Zhang, Vikram Voleti, Yuqun Wu, Mallikarjun B R, Shimon Vainer, Mark Boss, Yaoyao Liu
4Director makes video world models controllable by moving explicit 3D meshes through time while preserving realistic, consistent appearances.
EMPIRIC: Experiment-Driven Learning of Residual World Models for Robot Planning
Yichao Liang, Amber Li, Dat Nguyen, Emily Bunnapradist, Michelangelo Naim, Sreela Kodali, Matteo Merler, Bowen Li, Kiran Gopinathan, Yiyun Liu, Nikhil Pimpalkhare, Joshua B. Tenenbaum, Adrian Weller, Zenna Tavares, Tom Silver, Kevin Ellis
EMPIRIC lets robots discover missing physics through targeted experiments and use the resulting interpretable world models to plan better.
RoboCoach: World Models as Active Coaches for Compositional Robot Skills
Jiajun Liu, Yifan Chen, Yichao Liu, Jiayi Zhang, Ruoqu Chen, Shaoxuan Xie, Guocai Yao, Mengdi Xu, Sen Cui, Changshui Zhang
RoboCoach uses imagined robot failures to decide what demonstrations to request next, making long-horizon manipulation skills improve more efficiently.