NTH

AlayaWorld: Interactive Long-Horizon World Modeling -- Full Technical Report

AuthorsAlayaWorld Team, Kaipeng Zhang, Chuanhao Li, Yifan Zhan, Yongtao Ge, Yuanyang Yin, Jiaming Tan, Kang He, Liaoyuan Fan, Mingliang Zhai, Ruicong Liu, Xiaojie Xu, Xuangeng Chu, Zhen Li, Zhengyuan Lin, Zhixiang Wang, Zian Meng, Zihui Gao

July 23, 2026 2 min read
Watch on YouTube
The one-line take

AlayaWorld is a large interactive video world model designed to generate persistent, explorable environments over long time horizons while responding quickly to user prompts and camera actions.

Key results

222147
Training corpus

Total clips used across seven real and synthetic sources.

4
Distilled sampling steps

Inference steps per chunk after distillation, down from approximately 30.

0.9492
Brightness Consistency

AlayaWorld score on iWorld-Bench.

0.7985
Trajectory Accuracy

AlayaWorld score on iWorld-Bench.

0.8871
Memory Symmetry

AlayaWorld score on iWorld-Bench.

What the paper found

AlayaWorld, from Alaya Lab, is an interactive long-horizon video world model built on a 15B-class video diffusion transformer derived from LTX-2.3. It generates 24-fps video at 540p and 720p, autoregressively predicting short latent chunks from camera trajectories and switchable text prompts. Its central contribution is a bounded visual context combining a persistent sink frame, compressed temporal history, geometry-aligned spatial memory rendered with Depth-Anything-3, and recent-frame conditioning; this keeps computation approximately constant as the rollout grows while supporting revisits and scene consistency. Training uses 222147 clips spanning real captures, gameplay, and synthetic events, then applies corrupted histories, Helios-style drift simulation, and residual replay from the model’s own rollouts to improve recovery from accumulated errors. A discrete distillation scheme combining Distribution-Matching Distillation, self-forcing++, and consistency distillation reduces sampling from approximately 30 steps to 4 steps per chunk. On iWorld-Bench, AlayaWorld reaches 0.9492 in Brightness Consistency, 0.7985 in Trajectory Accuracy, and 0.8871 in Memory Symmetry, leading most reported metrics against systems including NVIDIA Cosmos, Tencent’s HunyuanVideo-1.5, Yume 1.5, Matrix-Game 2.0, and HY-World 1.5. Gemini and Kimi-K2.6 are cited as alternative and default vision-language annotation backends, respectively. The open-source system demonstrates controllable navigation, prompt-driven events, and stable leave-and-return generation, while remaining limited in explicit physical causality and object-state reasoning.

Original abstract

Unlike conventional video game development, which relies on labor-intensive pipelines for asset production, animation, physics, and programming, video world models generate interactive environments from user inputs instantly. It enable us to create customized, explorable, and continuously evolving virtual world from text, an image, or video. Realizing this vision requires four tightly coupled capabilities: interaction, persistent spatiotemporal consistency, stable long-horizon generation, and efficient response. We present AlayaWorld, an interactive long-horizon video world model that generates 24-fps video at 540p and 720p. Built on a 15B video diffusion transformer, AlayaWorld generates short latent chunks autoregressively under camera trajectories and switchable text prompts. Its bounded visual context combines a persistent sink frame, compressed temporal history, geometry-aligned spatial memory, and recent-frame conditioning. To reduce long-term drift, the model is trained with corrupted histories and prediction residuals collected from its own roll-outs. We further introduce a discrete autoregressive distillation formulation that combines distribution-matching distillation, self-forcing++, and consistency distillation, reducing inference from approximately 30 sampling steps to four steps per chunk. On iWorld-Bench, AlayaWorld achieves the best performance over long-horizon generation. Conceived as a full-stack, open-source, and long-term project, AlayaWorld is intended to provide an extensible foundation for future research on interactive video world models.

Read the original paper

More in World Models

Browse all 41 papers →
01World Model

4Director: Controlling Video World Models with Rigid 3D Geometry

Wei Cao, Hao Zhang, Vikram Voleti, Yuqun Wu, Mallikarjun B R, Shimon Vainer, Mark Boss, Yaoyao Liu

4Director makes video world models controllable by moving explicit 3D meshes through time while preserving realistic, consistent appearances.

Read analysis
02World Model

EMPIRIC: Experiment-Driven Learning of Residual World Models for Robot Planning

Yichao Liang, Amber Li, Dat Nguyen, Emily Bunnapradist, Michelangelo Naim, Sreela Kodali, Matteo Merler, Bowen Li, Kiran Gopinathan, Yiyun Liu, Nikhil Pimpalkhare, Joshua B. Tenenbaum, Adrian Weller, Zenna Tavares, Tom Silver, Kevin Ellis

EMPIRIC lets robots discover missing physics through targeted experiments and use the resulting interpretable world models to plan better.

Read analysis
03World Model

RoboCoach: World Models as Active Coaches for Compositional Robot Skills

Jiajun Liu, Yifan Chen, Yichao Liu, Jiayi Zhang, Ruoqu Chen, Shaoxuan Xie, Guocai Yao, Mengdi Xu, Sen Cui, Changshui Zhang

RoboCoach uses imagined robot failures to decide what demonstrations to request next, making long-horizon manipulation skills improve more efficiently.

Read analysis