AlayaWorld: Interactive Long-Horizon World Modeling -- Full Technical Report
AuthorsAlayaWorld Team, Kaipeng Zhang, Chuanhao Li, Yifan Zhan, Yongtao Ge, Yuanyang Yin, Jiaming Tan, Kang He, Liaoyuan Fan, Mingliang Zhai, Ruicong Liu, Xiaojie Xu, Xuangeng Chu, Zhen Li, Zhengyuan Lin, Zhixiang Wang, Zian Meng, Zihui Gao
Resources
AlayaWorld is a large interactive video world model designed to generate persistent, explorable environments over long time horizons while responding quickly to user prompts and camera actions.
Key results
Total clips used across seven real and synthetic sources.
Inference steps per chunk after distillation, down from approximately 30.
AlayaWorld score on iWorld-Bench.
AlayaWorld score on iWorld-Bench.
AlayaWorld score on iWorld-Bench.
What the paper found
AlayaWorld, from Alaya Lab, is an interactive long-horizon video world model built on a 15B-class video diffusion transformer derived from LTX-2.3. It generates 24-fps video at 540p and 720p, autoregressively predicting short latent chunks from camera trajectories and switchable text prompts. Its central contribution is a bounded visual context combining a persistent sink frame, compressed temporal history, geometry-aligned spatial memory rendered with Depth-Anything-3, and recent-frame conditioning; this keeps computation approximately constant as the rollout grows while supporting revisits and scene consistency. Training uses 222147 clips spanning real captures, gameplay, and synthetic events, then applies corrupted histories, Helios-style drift simulation, and residual replay from the model’s own rollouts to improve recovery from accumulated errors. A discrete distillation scheme combining Distribution-Matching Distillation, self-forcing++, and consistency distillation reduces sampling from approximately 30 steps to 4 steps per chunk. On iWorld-Bench, AlayaWorld reaches 0.9492 in Brightness Consistency, 0.7985 in Trajectory Accuracy, and 0.8871 in Memory Symmetry, leading most reported metrics against systems including NVIDIA Cosmos, Tencent’s HunyuanVideo-1.5, Yume 1.5, Matrix-Game 2.0, and HY-World 1.5. Gemini and Kimi-K2.6 are cited as alternative and default vision-language annotation backends, respectively. The open-source system demonstrates controllable navigation, prompt-driven events, and stable leave-and-return generation, while remaining limited in explicit physical causality and object-state reasoning.
Original abstract
Unlike conventional video game development, which relies on labor-intensive pipelines for asset production, animation, physics, and programming, video world models generate interactive environments from user inputs instantly. It enable us to create customized, explorable, and continuously evolving virtual world from text, an image, or video. Realizing this vision requires four tightly coupled capabilities: interaction, persistent spatiotemporal consistency, stable long-horizon generation, and efficient response. We present AlayaWorld, an interactive long-horizon video world model that generates 24-fps video at 540p and 720p. Built on a 15B video diffusion transformer, AlayaWorld generates short latent chunks autoregressively under camera trajectories and switchable text prompts. Its bounded visual context combines a persistent sink frame, compressed temporal history, geometry-aligned spatial memory, and recent-frame conditioning. To reduce long-term drift, the model is trained with corrupted histories and prediction residuals collected from its own roll-outs. We further introduce a discrete autoregressive distillation formulation that combines distribution-matching distillation, self-forcing++, and consistency distillation, reducing inference from approximately 30 sampling steps to four steps per chunk. On iWorld-Bench, AlayaWorld achieves the best performance over long-horizon generation. Conceived as a full-stack, open-source, and long-term project, AlayaWorld is intended to provide an extensible foundation for future research on interactive video world models.
Read the original paperMore in World Models
Browse all 41 papers →4Director: Controlling Video World Models with Rigid 3D Geometry
Wei Cao, Hao Zhang, Vikram Voleti, Yuqun Wu, Mallikarjun B R, Shimon Vainer, Mark Boss, Yaoyao Liu
4Director makes video world models controllable by moving explicit 3D meshes through time while preserving realistic, consistent appearances.
EMPIRIC: Experiment-Driven Learning of Residual World Models for Robot Planning
Yichao Liang, Amber Li, Dat Nguyen, Emily Bunnapradist, Michelangelo Naim, Sreela Kodali, Matteo Merler, Bowen Li, Kiran Gopinathan, Yiyun Liu, Nikhil Pimpalkhare, Joshua B. Tenenbaum, Adrian Weller, Zenna Tavares, Tom Silver, Kevin Ellis
EMPIRIC lets robots discover missing physics through targeted experiments and use the resulting interpretable world models to plan better.
RoboCoach: World Models as Active Coaches for Compositional Robot Skills
Jiajun Liu, Yifan Chen, Yichao Liu, Jiayi Zhang, Ruoqu Chen, Shaoxuan Xie, Guocai Yao, Mengdi Xu, Sen Cui, Changshui Zhang
RoboCoach uses imagined robot failures to decide what demonstrations to request next, making long-horizon manipulation skills improve more efficiently.