Flex-$π$: A Multi-Stream World-Action Model with Compute Flexibility
AuthorsGe Yan, Jinghao Liu, Yuzhi Fan, Lei Cai, Minwen Liao, Jesse Zhang, Dieter Fox
Resources
Flex-π gives robots a flexible world model that combines video, 3D structure, semantics, and actions to improve precise bimanual manipulation while adapting compute at inference time.
Key results
FLEX-π parameter count
Hours from AGIBOT World
50-task average for action-only FLEX-π
Fixed full-joint FLEX-π variant
Milliseconds per call on the real-robot deployment stack
Average RoboTwin success reduction without pointmap training
What the paper found
FLEX-π is a 6B-parameter world-action model that jointly predicts robot actions with future RGB, 3D pointmaps, and object-centric DINOv3 features. Its key finding is that the frozen Wan-2.2 VAE, trained for RGB video, reconstructs pointmaps accurately enough to provide geometric supervision without new sensors or VAE pre-training, while Depth Anything 3 supplies pointmaps from RGB and DINOv3 supplies semantic tokens. A Mixture-of-Transformers backbone fuses these streams, and independent 0.5-probability visual dropout with cross-modality forcing trains one checkpoint to operate from any subset of inputs and generate any subset of futures, from action-only to full joint inference. Pre-trained on 500 hours of AGIBOT World, FLEX-π reaches 94.6% randomized success on the 50-task RoboTwin benchmark, and its fixed full-joint variant reaches 99.2% on LIBERO, matching Qwen-RobotManip while using far less robot data. On a real bimanual YAM robot, it outperforms π0.5, ManiFlow, and Fast-WAM on five dexterous tasks; action-only inference runs at 60 ms per call and already exceeds the baselines, while joint visual generation raises average success by 13%. The additional geometry stream is consequential: removing pointmaps reduces RoboTwin success by 20%, and cross-modality forcing improves representation learning even when modalities are available at deployment.
Original abstract
World-action models (WAMs) predict the future to act better, but nearly all of them predict only RGB latents, trained purely for pixel reconstruction, with no explicit signal for the 3D geometry or object semantics manipulation needs. We find a surprising free lunch: the same frozen video-generation VAE that encodes RGB also encodes 3D pointmaps almost losslessly, with no pointmap-specific training at all. This lets us supervise Flex-$π$, a 6B-parameter WAM, on 3D geometry and object-centric DINO semantics alongside RGB, at no cost in new sensors, new pre-training, or inference latency. Every visual signal is projected into this shared latent space and denoised jointly with actions inside a Mixture-of-Transformers backbone; per-stream dropout with cross-modality forcing then lets a single trained checkpoint run on any subset of these streams, from a fast action-only mode to full joint generation. The result is a policy that is exceptionally demonstration-efficient and generalizes well, beating the strongest baselines by up to 2-7$\times$ on dexterous, precise, real-world bimanual manipulation tasks both in and out of distribution, all while running faster than $π_{0.5}$. Our project website: https://flex-pi.github.io/
Read the original paperMore in World Models
Browse all 41 papers →4Director: Controlling Video World Models with Rigid 3D Geometry
Wei Cao, Hao Zhang, Vikram Voleti, Yuqun Wu, Mallikarjun B R, Shimon Vainer, Mark Boss, Yaoyao Liu
4Director makes video world models controllable by moving explicit 3D meshes through time while preserving realistic, consistent appearances.
EMPIRIC: Experiment-Driven Learning of Residual World Models for Robot Planning
Yichao Liang, Amber Li, Dat Nguyen, Emily Bunnapradist, Michelangelo Naim, Sreela Kodali, Matteo Merler, Bowen Li, Kiran Gopinathan, Yiyun Liu, Nikhil Pimpalkhare, Joshua B. Tenenbaum, Adrian Weller, Zenna Tavares, Tom Silver, Kevin Ellis
EMPIRIC lets robots discover missing physics through targeted experiments and use the resulting interpretable world models to plan better.
RoboCoach: World Models as Active Coaches for Compositional Robot Skills
Jiajun Liu, Yifan Chen, Yichao Liu, Jiayi Zhang, Ruoqu Chen, Shaoxuan Xie, Guocai Yao, Mengdi Xu, Sen Cui, Changshui Zhang
RoboCoach uses imagined robot failures to decide what demonstrations to request next, making long-horizon manipulation skills improve more efficiently.