World-Language-Action Model for Unified World Modeling, Language Reasoning, and Action Synthesis
AuthorsYi Yang, Zhihong Liu, Siqi Kou, Yiyang Chen, Yanzhe Hu, Jianbo Zhou, Boyuan Zhao, Zhijie Wei, Xiao Xia, Xueqi Li, Pengfei Liu, Zhijie Deng
Resources
This paper introduces a robot foundation model that can predict what happens next, reason in language, and generate actions—all in one system for better long-horizon embodied tasks.
Key results
WLA-0 active parameters during inference
Inference latency on an NVIDIA RTX 5090 in ms
Best clean-scene success rate reported for WLA-0
Average success rate on the memory-dependent benchmark
Average success rate across Spatial, Object, Goal, and Long suites
Average success rate with test-time scaling
What the paper found
World Language Action Model, or WLA, is a new embodied foundation model from SJTU and collaborators that unifies world modeling, language reasoning, and robot action synthesis by predicting a semantic textual intention, a future visual state, and executable action chunks in one autoregressive Transformer. Unlike diffusion-based World Action Models and typical Vision-Language-Action systems, WLA uses meta-queries plus a dedicated World Expert to separate compact latent dynamics from fine-grained frame prediction, then lets an Action Expert condition on that latent transition; this makes the world module removable at inference while still supporting test-time scaling through imagined rollouts and value-based trajectory selection. The WLA-0 prototype uses 2B active parameters for inference, reaches about 40 ms latency on an NVIDIA RTX 5090, and removes the need for embodied pretraining. On RoboTwin 2.0, WLA-0 scores 92.94% on clean scenes, and on RMBench it reaches 56.5% average success, nearly doubling the strongest prior memory-based baseline; on LIBERO it achieves 98.6% average success, rising to 98.9% with test-time scaling. The paper’s key claim is that textual subtasks are not just instruction paraphrases but the control interface for long-horizon progress tracking, memory, and recovery, enabling cross-embodiment video-only task acquisition without action annotations.
Original abstract
We propose world-language-action (WLA) models as a new class of embodied foundation models. WLA takes textual instructions, images, and robot states as inputs to jointly predict textual subtasks, subgoal images, and robot actions, conjoining the \emph{world modeling interface} to learn from extensive egocentric videos as in the world-action model (WAM) and the \emph{language reasoning} capacities to solve complex long-horizon tasks as in vision-language-action (VLA) models. At the core of WLA lies an \emph{autoregressive (AR)} Transformer backbone, instead of a bidirectional diffusion Transformer as in WAMs, to predict the \emph{next state}, comprising the \emph{semantic-level} textual intention and complementary \emph{fine-grained} physical dynamics. The physical dynamics are supervised by the world modeling objective based on a dedicated World Expert, and are leveraged to ease the characterization of the state-action correlation for the Action Expert. WLA leverages meta-queries to make the world prediction \emph{implicitly} impact the action generation so that the former can be disabled during inference. The world prediction can also be activated to enable test-time scaling for improved robot control. Our WLA-0 prototype, with 2B active parameters, achieves 40 ms per inference on an NVIDIA RTX 5090. Evaluations across simulated and real-world environments demonstrate that WLA-0 achieves state-of-the-art multi-task and long-horizon learning abilities, e.g., 92.94\% success rate on RoboTwin2.0 Clean and 56.5\% success rate on RMBench. WLA-0 also holds the promise to learn novel tasks directly from \emph{cross-embodiment robot videos} without action annotations.
Read the original paperMore in Embodied AI
Browse all 48 papers →GroundingPI: A Grounding Foundation Model towards Physical Intelligence with Visual Primitives
Qize Yu, Lianrui Fan, Boyu Chen, Jiaqi Liang, Xini Ding, Yue Chen, Zetian Song, Yuran Wang, Yi Zou, Kaixuan Wang, Tianxing Chen, Wenxuan Song, Bohan Zhou, Mingleyang Li, Siqiao Huang, Yuqi Ye, Caigao Jiang, Wei Wei, Ruihai Wu, Hang Zhang, Yixiao Ge, Shuchang Zhou, Shilong Liu, Xianming Liu, Ping Luo, Shiyu Huang
GroundingPI argues that fast, precise visual grounding should be the perceptual foundation for capable robots and autonomous vehicles.
MM-ABC: Towards Generalist Mobile Manipulation via Seeing, Coordinating and Imagining
Qiwei Liang, Guangyu Chen, Shaolong Zhu, Zikuan Xiao, Jinxuan Lu, Yifan Xie, Renjing Xu, Wenbo Ding, Tianxing Chen
MM-ABC is a generalist robot foundation model that helps mobile manipulators see their surroundings, coordinate arm and base motion, and imagine future actions for better performance.
Morphometric Imitation: From Morphology and Contact Aware Hand Retargeting to Sim-to-Real Visuomotor Policy
Tara Sadjadpour, Siming He, C. K. Wolfe, Haozhi Qi, Lea Wilken, S. Shankar Sastry, Claire Tomlin, Jitendra Malik
A three-stage system converts human hand demonstrations into robust, zero-shot real-robot dexterous manipulation policies across different hand morphologies.