Programmable World Model
AuthorsZheng-Hui Huang, Guixu Lin, Jiacheng Lin, Yi-Chuan Huang, Ruihan Yu, Muyao Niu, Siqi Yang, Yu-Lun Liu, Yung-Yu Chuang, Kaipeng Zhang, Zhixiang Wang
Resources
A programmable world model combines an explicit game-state engine with generative video to create persistent, controllable interactive worlds.
Key results
Number of benchmark clips used to evaluate world-state consistency.
Accuracy of visible alive-character counts on CombatStateBench.
Accuracy for visually realizing engine-recorded death states.
VBench temporal-stability score achieved by the proposed renderer.
What the paper found
Programmable World Model addresses a core weakness of interactive video generators: they can produce plausible frames but do not reliably remember what remains true across long interactions. Instead of asking a video model to infer game logic, the system uses a VLM-based coding agent to convert natural-language instructions into executable rules, while a lightweight engine maintains persistent state for entities, including off-screen objects, health, inventory, identity, and event history. State-augmented 3D oriented bounding boxes provide a compact world-space representation, and a deterministic compiler projects them into identity, semantic, and seven-way motion controls aligned with the target camera. These controls condition a pretrained LingBot-World-v1 renderer through a trainable Structured Spatial ControlNet, with temporal history and geometry-aligned spatial memory supporting chunk-autoregressive rollouts. Training annotations are automatically recovered from gameplay videos from Cyberpunk 2077, Forza Horizon 6, and Grand Theft Auto V using ViPE, Qwen3-VL, SAM3, and WildDet3D. On the 50-clip CombatStateBench, the method reaches 94% Count Accuracy and 98% State Accuracy, compared with 40.75% and 8.00% for LingBot-World-V2; it also achieves 99.00% Temporal Stability. The result is a programmable, playable generative world where state evolution is explicit and verifiable, while visual appearance remains the responsibility of the generative renderer.
Original abstract
Recent video world models generate increasingly realistic and interactive visual experiences, yet lack reliable mechanisms for maintaining persistent world state and enforcing programmable rules over extended interactions. We introduce Programmable World Model, a framework that decouples world-state evolution from visual observation generation. An agent translates natural-language instructions into executable programs that specify entity states and state-transition rules, enabling direct control over individual entities and their interactions. A lightweight engine executes these programs to update and maintain an explicit, persistent global world state, including off-screen entities and non-visual attributes. To connect world state with visual generation, we introduce state-augmented 3D oriented bounding boxes (OBBs) as an intermediate representation. This representation, together with the target camera trajectory, is deterministically compiled into pixel-aligned spatiotemporal conditioning signals for a pretrained video model serving as the generative renderer. This design allows users to create playable games with predefined mechanics, direct control over individual entities, and persistent world state throughout gameplay. We further introduce CombatStateBench, a benchmark for evaluating programmable world models. On CombatStateBench, our method achieves 94% Count Accuracy and 98% State Accuracy, substantially outperforming existing interactive video world models while supporting coherent long-horizon generation. These results demonstrate the effectiveness of separating explicit state evolution from generative rendering for building persistent, programmable worlds.
Read the original paperMore in World Models
Browse all 41 papers →4Director: Controlling Video World Models with Rigid 3D Geometry
Wei Cao, Hao Zhang, Vikram Voleti, Yuqun Wu, Mallikarjun B R, Shimon Vainer, Mark Boss, Yaoyao Liu
4Director makes video world models controllable by moving explicit 3D meshes through time while preserving realistic, consistent appearances.
EMPIRIC: Experiment-Driven Learning of Residual World Models for Robot Planning
Yichao Liang, Amber Li, Dat Nguyen, Emily Bunnapradist, Michelangelo Naim, Sreela Kodali, Matteo Merler, Bowen Li, Kiran Gopinathan, Yiyun Liu, Nikhil Pimpalkhare, Joshua B. Tenenbaum, Adrian Weller, Zenna Tavares, Tom Silver, Kevin Ellis
EMPIRIC lets robots discover missing physics through targeted experiments and use the resulting interpretable world models to plan better.
RoboCoach: World Models as Active Coaches for Compositional Robot Skills
Jiajun Liu, Yifan Chen, Yichao Liu, Jiayi Zhang, Ruoqu Chen, Shaoxuan Xie, Guocai Yao, Mengdi Xu, Sen Cui, Changshui Zhang
RoboCoach uses imagined robot failures to decide what demonstrations to request next, making long-horizon manipulation skills improve more efficiently.