NTH

From Pixels to States: Rethinking Interactive World Models as Game Engines

AuthorsZhen Li, Zian Meng, Shuwei Shi, Mingliang Zhai, Jiaming Tan, Chuanhao Li, Kaipeng Zhang

July 23, 2026 2 min read
Watch on YouTube
The one-line take

This paper argues that truly interactive AI game worlds need explicit, persistent state dynamics and introduces a large-scale dataset and framework to move beyond pixel-only video prediction.

Key results

90h
Black Myth: Wukong gameplay

The proposed state-aware dataset contains more than 90h of boss-encounter gameplay.

1280x720
Video resolution

Gameplay was recorded at 1280 × 720 resolution.

30
Capture frame rate

The dataset was captured at 30 FPS.

235B
Semantic captioning model

Semantic annotations use Qwen3-VL-235B-A22B-Instruct, identified by its 235B model scale.

What the paper found

“From Pixels to States” from Alaya Lab argues that interactive world models should be evaluated as game engines, not merely video predictors. The paper organizes the field around the conventional action-state-observation loop and identifies four requirements: player action control, game-state dynamics, state-observation persistence, and real-time generation. It compares geometric camera control, device-level motor signals, and semantic event interfaces; contrasts pixel-entangled, learned-latent, and explicit symbolic states; and separates memory that retrieves past observations from memory that estimates the world’s current condition. The central diagnosis is that systems such as OpenAI’s Sora-oriented world-simulation line, Decart’s Oasis, Google DeepMind’s Genie 3, Matrix-Game, and Hunyuan-GameCraft increasingly support exploration and low-latency interaction, but usually leave health, cooldowns, irreversible consequences, and outcome timing implicit in visual dynamics. To address the data bottleneck, the authors build a scalable Black Myth: Wukong data engine producing more than 90h of boss gameplay at 1280 × 720 and 30 FPS, with frame-aligned keyboard and mouse actions, engine-exported states, RGB frames, depth maps, and structured and semantic captions. Semantic annotations are generated with Qwen3-VL-235B-A22B-Instruct. The paper’s main research direction is to make explicit state transitions drive generation, ground memory updates in those transitions, and distinguish control latency from the rule-accurate timing of consequences.

Original abstract

Building interactive worlds that respond coherently to player actions has long been a shared goal of computer graphics, games, and artificial intelligence. Recent video generative models provide a data-driven route toward this goal by predicting future observations conditioned on user actions, and are increasingly regarded as potential next-generation game engines. Realizing a genuinely interactive game world, however, requires interaction outcomes that follow rules over evolving game conditions, consequences that persist over long horizons, and a generation loop that operates in real time. Conventional game engines realize these properties through a recurrent action-state-observation loop, in which player actions update an explicit game state according to predefined rules and observations are rendered from the resulting state. Taking this loop as an organizing lens, this paper examines interactive game world modeling along four dimensions: player action control, game state dynamics, state-observation persistence, and real-time interactive generation. For each dimension, we start from the capabilities required by an interactive game world, group existing approaches into representative families, and discuss the strengths and trade-offs of each family. Complementing this analysis, we present a scalable data engine for Black Myth: Wukong that collects over 90 hours of gameplay with frame-aligned player actions, ground-truth game states, and visual observations, together with structured and semantic annotations, as a resource for state-aware game world modeling. We hope this paper offers a clear picture of where the field stands and fosters progress toward interactive game worlds.

Read the original paper

More in World Models

Browse all 41 papers →
01World Model

4Director: Controlling Video World Models with Rigid 3D Geometry

Wei Cao, Hao Zhang, Vikram Voleti, Yuqun Wu, Mallikarjun B R, Shimon Vainer, Mark Boss, Yaoyao Liu

4Director makes video world models controllable by moving explicit 3D meshes through time while preserving realistic, consistent appearances.

Read analysis
02World Model

EMPIRIC: Experiment-Driven Learning of Residual World Models for Robot Planning

Yichao Liang, Amber Li, Dat Nguyen, Emily Bunnapradist, Michelangelo Naim, Sreela Kodali, Matteo Merler, Bowen Li, Kiran Gopinathan, Yiyun Liu, Nikhil Pimpalkhare, Joshua B. Tenenbaum, Adrian Weller, Zenna Tavares, Tom Silver, Kevin Ellis

EMPIRIC lets robots discover missing physics through targeted experiments and use the resulting interpretable world models to plan better.

Read analysis
03World Model

RoboCoach: World Models as Active Coaches for Compositional Robot Skills

Jiajun Liu, Yifan Chen, Yichao Liu, Jiayi Zhang, Ruoqu Chen, Shaoxuan Xie, Guocai Yao, Mengdi Xu, Sen Cui, Changshui Zhang

RoboCoach uses imagined robot failures to decide what demonstrations to request next, making long-horizon manipulation skills improve more efficiently.

Read analysis