NTH

Programmable World Model

AuthorsZheng-Hui Huang, Guixu Lin, Jiacheng Lin, Yi-Chuan Huang, Ruihan Yu, Muyao Niu, Siqi Yang, Yu-Lun Liu, Yung-Yu Chuang, Kaipeng Zhang, Zhixiang Wang

September 15, 2026 2 min read
Watch on YouTube
The one-line take

A programmable world model combines an explicit game-state engine with generative video to create persistent, controllable interactive worlds.

Key results

50
CombatStateBench size

Number of benchmark clips used to evaluate world-state consistency.

94%
Count Accuracy

Accuracy of visible alive-character counts on CombatStateBench.

98%
State Accuracy

Accuracy for visually realizing engine-recorded death states.

99.00%
Temporal Stability

VBench temporal-stability score achieved by the proposed renderer.

What the paper found

Programmable World Model addresses a core weakness of interactive video generators: they can produce plausible frames but do not reliably remember what remains true across long interactions. Instead of asking a video model to infer game logic, the system uses a VLM-based coding agent to convert natural-language instructions into executable rules, while a lightweight engine maintains persistent state for entities, including off-screen objects, health, inventory, identity, and event history. State-augmented 3D oriented bounding boxes provide a compact world-space representation, and a deterministic compiler projects them into identity, semantic, and seven-way motion controls aligned with the target camera. These controls condition a pretrained LingBot-World-v1 renderer through a trainable Structured Spatial ControlNet, with temporal history and geometry-aligned spatial memory supporting chunk-autoregressive rollouts. Training annotations are automatically recovered from gameplay videos from Cyberpunk 2077, Forza Horizon 6, and Grand Theft Auto V using ViPE, Qwen3-VL, SAM3, and WildDet3D. On the 50-clip CombatStateBench, the method reaches 94% Count Accuracy and 98% State Accuracy, compared with 40.75% and 8.00% for LingBot-World-V2; it also achieves 99.00% Temporal Stability. The result is a programmable, playable generative world where state evolution is explicit and verifiable, while visual appearance remains the responsibility of the generative renderer.

Original abstract

Recent video world models generate increasingly realistic and interactive visual experiences, yet lack reliable mechanisms for maintaining persistent world state and enforcing programmable rules over extended interactions. We introduce Programmable World Model, a framework that decouples world-state evolution from visual observation generation. An agent translates natural-language instructions into executable programs that specify entity states and state-transition rules, enabling direct control over individual entities and their interactions. A lightweight engine executes these programs to update and maintain an explicit, persistent global world state, including off-screen entities and non-visual attributes. To connect world state with visual generation, we introduce state-augmented 3D oriented bounding boxes (OBBs) as an intermediate representation. This representation, together with the target camera trajectory, is deterministically compiled into pixel-aligned spatiotemporal conditioning signals for a pretrained video model serving as the generative renderer. This design allows users to create playable games with predefined mechanics, direct control over individual entities, and persistent world state throughout gameplay. We further introduce CombatStateBench, a benchmark for evaluating programmable world models. On CombatStateBench, our method achieves 94% Count Accuracy and 98% State Accuracy, substantially outperforming existing interactive video world models while supporting coherent long-horizon generation. These results demonstrate the effectiveness of separating explicit state evolution from generative rendering for building persistent, programmable worlds.

Read the original paper

More in World Models

Browse all 41 papers →
01World Model

4Director: Controlling Video World Models with Rigid 3D Geometry

Wei Cao, Hao Zhang, Vikram Voleti, Yuqun Wu, Mallikarjun B R, Shimon Vainer, Mark Boss, Yaoyao Liu

4Director makes video world models controllable by moving explicit 3D meshes through time while preserving realistic, consistent appearances.

Read analysis
02World Model

EMPIRIC: Experiment-Driven Learning of Residual World Models for Robot Planning

Yichao Liang, Amber Li, Dat Nguyen, Emily Bunnapradist, Michelangelo Naim, Sreela Kodali, Matteo Merler, Bowen Li, Kiran Gopinathan, Yiyun Liu, Nikhil Pimpalkhare, Joshua B. Tenenbaum, Adrian Weller, Zenna Tavares, Tom Silver, Kevin Ellis

EMPIRIC lets robots discover missing physics through targeted experiments and use the resulting interpretable world models to plan better.

Read analysis
03World Model

RoboCoach: World Models as Active Coaches for Compositional Robot Skills

Jiajun Liu, Yifan Chen, Yichao Liu, Jiayi Zhang, Ruoqu Chen, Shaoxuan Xie, Guocai Yao, Mengdi Xu, Sen Cui, Changshui Zhang

RoboCoach uses imagined robot failures to decide what demonstrations to request next, making long-horizon manipulation skills improve more efficiently.

Read analysis