NTH

Flex-$π$: A Multi-Stream World-Action Model with Compute Flexibility

AuthorsGe Yan, Jinghao Liu, Yuzhi Fan, Lei Cai, Minwen Liao, Jesse Zhang, Dieter Fox

August 21, 2026 2 min read
Watch on YouTube
The one-line take

Flex-π gives robots a flexible world model that combines video, 3D structure, semantics, and actions to improve precise bimanual manipulation while adapting compute at inference time.

Key results

6B
Model size

FLEX-π parameter count

500
Pre-training data

Hours from AGIBOT World

94.6%
RoboTwin randomized success

50-task average for action-only FLEX-π

99.2%
LIBERO success

Fixed full-joint FLEX-π variant

60
Action-only latency

Milliseconds per call on the real-robot deployment stack

20%
Pointmap ablation drop

Average RoboTwin success reduction without pointmap training

What the paper found

FLEX-π is a 6B-parameter world-action model that jointly predicts robot actions with future RGB, 3D pointmaps, and object-centric DINOv3 features. Its key finding is that the frozen Wan-2.2 VAE, trained for RGB video, reconstructs pointmaps accurately enough to provide geometric supervision without new sensors or VAE pre-training, while Depth Anything 3 supplies pointmaps from RGB and DINOv3 supplies semantic tokens. A Mixture-of-Transformers backbone fuses these streams, and independent 0.5-probability visual dropout with cross-modality forcing trains one checkpoint to operate from any subset of inputs and generate any subset of futures, from action-only to full joint inference. Pre-trained on 500 hours of AGIBOT World, FLEX-π reaches 94.6% randomized success on the 50-task RoboTwin benchmark, and its fixed full-joint variant reaches 99.2% on LIBERO, matching Qwen-RobotManip while using far less robot data. On a real bimanual YAM robot, it outperforms π0.5, ManiFlow, and Fast-WAM on five dexterous tasks; action-only inference runs at 60 ms per call and already exceeds the baselines, while joint visual generation raises average success by 13%. The additional geometry stream is consequential: removing pointmaps reduces RoboTwin success by 20%, and cross-modality forcing improves representation learning even when modalities are available at deployment.

Original abstract

World-action models (WAMs) predict the future to act better, but nearly all of them predict only RGB latents, trained purely for pixel reconstruction, with no explicit signal for the 3D geometry or object semantics manipulation needs. We find a surprising free lunch: the same frozen video-generation VAE that encodes RGB also encodes 3D pointmaps almost losslessly, with no pointmap-specific training at all. This lets us supervise Flex-$π$, a 6B-parameter WAM, on 3D geometry and object-centric DINO semantics alongside RGB, at no cost in new sensors, new pre-training, or inference latency. Every visual signal is projected into this shared latent space and denoised jointly with actions inside a Mixture-of-Transformers backbone; per-stream dropout with cross-modality forcing then lets a single trained checkpoint run on any subset of these streams, from a fast action-only mode to full joint generation. The result is a policy that is exceptionally demonstration-efficient and generalizes well, beating the strongest baselines by up to 2-7$\times$ on dexterous, precise, real-world bimanual manipulation tasks both in and out of distribution, all while running faster than $π_{0.5}$. Our project website: https://flex-pi.github.io/

Read the original paper

More in World Models

Browse all 41 papers →
01World Model

4Director: Controlling Video World Models with Rigid 3D Geometry

Wei Cao, Hao Zhang, Vikram Voleti, Yuqun Wu, Mallikarjun B R, Shimon Vainer, Mark Boss, Yaoyao Liu

4Director makes video world models controllable by moving explicit 3D meshes through time while preserving realistic, consistent appearances.

Read analysis
02World Model

EMPIRIC: Experiment-Driven Learning of Residual World Models for Robot Planning

Yichao Liang, Amber Li, Dat Nguyen, Emily Bunnapradist, Michelangelo Naim, Sreela Kodali, Matteo Merler, Bowen Li, Kiran Gopinathan, Yiyun Liu, Nikhil Pimpalkhare, Joshua B. Tenenbaum, Adrian Weller, Zenna Tavares, Tom Silver, Kevin Ellis

EMPIRIC lets robots discover missing physics through targeted experiments and use the resulting interpretable world models to plan better.

Read analysis
03World Model

RoboCoach: World Models as Active Coaches for Compositional Robot Skills

Jiajun Liu, Yifan Chen, Yichao Liu, Jiayi Zhang, Ruoqu Chen, Shaoxuan Xie, Guocai Yao, Mengdi Xu, Sen Cui, Changshui Zhang

RoboCoach uses imagined robot failures to decide what demonstrations to request next, making long-horizon manipulation skills improve more efficiently.

Read analysis