NTH

World-Language-Action Model for Unified World Modeling, Language Reasoning, and Action Synthesis

AuthorsYi Yang, Zhihong Liu, Siqi Kou, Yiyang Chen, Yanzhe Hu, Jianbo Zhou, Boyuan Zhao, Zhijie Wei, Xiao Xia, Xueqi Li, Pengfei Liu, Zhijie Deng

June 15, 2026 2 min read
Watch on YouTube
The one-line take

This paper introduces a robot foundation model that can predict what happens next, reason in language, and generate actions—all in one system for better long-horizon embodied tasks.

Key results

2B
active parameters

WLA-0 active parameters during inference

40
latency

Inference latency on an NVIDIA RTX 5090 in ms

92.94%
RoboTwin 2.0 Clean success

Best clean-scene success rate reported for WLA-0

56.5%
RMBench average success

Average success rate on the memory-dependent benchmark

98.6%
LIBERO average success

Average success rate across Spatial, Object, Goal, and Long suites

98.9%
LIBERO +TTS average success

Average success rate with test-time scaling

What the paper found

World Language Action Model, or WLA, is a new embodied foundation model from SJTU and collaborators that unifies world modeling, language reasoning, and robot action synthesis by predicting a semantic textual intention, a future visual state, and executable action chunks in one autoregressive Transformer. Unlike diffusion-based World Action Models and typical Vision-Language-Action systems, WLA uses meta-queries plus a dedicated World Expert to separate compact latent dynamics from fine-grained frame prediction, then lets an Action Expert condition on that latent transition; this makes the world module removable at inference while still supporting test-time scaling through imagined rollouts and value-based trajectory selection. The WLA-0 prototype uses 2B active parameters for inference, reaches about 40 ms latency on an NVIDIA RTX 5090, and removes the need for embodied pretraining. On RoboTwin 2.0, WLA-0 scores 92.94% on clean scenes, and on RMBench it reaches 56.5% average success, nearly doubling the strongest prior memory-based baseline; on LIBERO it achieves 98.6% average success, rising to 98.9% with test-time scaling. The paper’s key claim is that textual subtasks are not just instruction paraphrases but the control interface for long-horizon progress tracking, memory, and recovery, enabling cross-embodiment video-only task acquisition without action annotations.

Original abstract

We propose world-language-action (WLA) models as a new class of embodied foundation models. WLA takes textual instructions, images, and robot states as inputs to jointly predict textual subtasks, subgoal images, and robot actions, conjoining the \emph{world modeling interface} to learn from extensive egocentric videos as in the world-action model (WAM) and the \emph{language reasoning} capacities to solve complex long-horizon tasks as in vision-language-action (VLA) models. At the core of WLA lies an \emph{autoregressive (AR)} Transformer backbone, instead of a bidirectional diffusion Transformer as in WAMs, to predict the \emph{next state}, comprising the \emph{semantic-level} textual intention and complementary \emph{fine-grained} physical dynamics. The physical dynamics are supervised by the world modeling objective based on a dedicated World Expert, and are leveraged to ease the characterization of the state-action correlation for the Action Expert. WLA leverages meta-queries to make the world prediction \emph{implicitly} impact the action generation so that the former can be disabled during inference. The world prediction can also be activated to enable test-time scaling for improved robot control. Our WLA-0 prototype, with 2B active parameters, achieves 40 ms per inference on an NVIDIA RTX 5090. Evaluations across simulated and real-world environments demonstrate that WLA-0 achieves state-of-the-art multi-task and long-horizon learning abilities, e.g., 92.94\% success rate on RoboTwin2.0 Clean and 56.5\% success rate on RMBench. WLA-0 also holds the promise to learn novel tasks directly from \emph{cross-embodiment robot videos} without action annotations.

Read the original paper

More in Embodied AI

Browse all 48 papers →
01Embodied Ai

GroundingPI: A Grounding Foundation Model towards Physical Intelligence with Visual Primitives

Qize Yu, Lianrui Fan, Boyu Chen, Jiaqi Liang, Xini Ding, Yue Chen, Zetian Song, Yuran Wang, Yi Zou, Kaixuan Wang, Tianxing Chen, Wenxuan Song, Bohan Zhou, Mingleyang Li, Siqiao Huang, Yuqi Ye, Caigao Jiang, Wei Wei, Ruihai Wu, Hang Zhang, Yixiao Ge, Shuchang Zhou, Shilong Liu, Xianming Liu, Ping Luo, Shiyu Huang

GroundingPI argues that fast, precise visual grounding should be the perceptual foundation for capable robots and autonomous vehicles.

Read analysis