SolarWM: Open Data and Scalable Training for Long-Horizon Video World Models
AuthorsJunchao Huang, Guian Fang, Shengju Qian, Xianghao Kong, Zhuoran Zhao, Wei Huang, Yihua Du, Zixin Zhang, Justin Cui, Yuchao Gu, Yukang Chen, Xinting Hu, Tianyu He, Shaoshuai Shi, Zhuotao Tian, Xin Wang, Mike Zheng Shou, Li Jiang
Resources
SolarWM is an open, scalable toolkit for training video world models that can simulate interactive environments over minutes or even hours.
Key results
Unified clips processed by the SolarWM open data engine.
Real-world, synthetic, and game datasets contributing to the corpus.
Physical storage size in TB for the complete corpus.
Parameter count of the SolarWM-minimax-h3-33B model.
Seconds used for training sequences.
Minutes reached by uninterrupted autoregressive generation.
What the paper found
SolarWM presents an open, reconfigurable foundation for interactive video world models, addressing the field’s central mismatch between heterogeneous video data and incompatible generation backbones. Its data engine converts 1.43M canonical clips from 10 datasets, spanning real, synthetic, and game environments, into frame-aligned records containing video, metric camera geometry, dense captions, quality scores, filtering decisions, and provenance, with the corpus occupying 25.85 TB. The framework supports four backbone-native models built on Wan2.2, LTX-2.5, and MiniMax-H3, including a largest 33B-parameter variant, while related open systems include Google’s Genie 3, Microsoft’s MineWorld, and NVIDIA’s SANA-WM. Training uses three stages: bidirectional camera-conditioned adaptation with fused-PRoPE, teacher-forced AnyFlow autoregressive initialization, and distribution-matching distillation, eliminating specialized ODE or consistency-distillation initialization. The causal models use four sampling steps at 16 fps and are trained only on 5s sequences, yet produce real-time, uninterrupted rollouts reaching a demonstrated 60-minute horizon without attention sinks, long-sequence fine-tuning, reference-frame resets, or clip splicing. The main contribution is not a single architecture, but a reproducible stack that separates expensive source preprocessing from configurable data mixtures while preserving each backbone’s native representation and objective.
Original abstract
We introduce SolarWM, a fully open foundation for building interactive video world models from data preparation through long-horizon inference. Training across heterogeneous data sources and video backbones is challenging: datasets differ in temporal scale, camera geometry, visual quality, motion, and captioning styles, while video generators use distinct representations and architectures. Naive data mixing and model-specific implementations therefore produce inconsistent supervision and make results difficult to reproduce and compare. SolarWM addresses this coupling with a reconfigurable multi-source data engine and a backbone-native adaptation framework. The engine converts 1.43 million canonical clips from 10 datasets into a unified, frame-aligned contract covering visual observations, metric camera geometry, captions, quality metadata, selection decisions, and provenance, while decoupling source processing from mixture construction. Under shared camera-conditioning, training, and inference interfaces, we instantiate four 5B--33B models based on Wan2.2, LTX-2.5, and MiniMax-H3 while preserving their native representations and objectives. A unified three-stage recipe combines bidirectional adaptation, teacher-forced autoregressive initialization, and distribution matching distillation. The resulting causal models enable real-time interaction over rollouts ranging from minutes to hours after being trained on only 5s sequences. By releasing the resulting data, pipeline, recipes, weights, and framework, SolarWM provides a reproducible and extensible foundation for interactive world-model research.
Read the original paperMore in World Models
Browse all 41 papers →4Director: Controlling Video World Models with Rigid 3D Geometry
Wei Cao, Hao Zhang, Vikram Voleti, Yuqun Wu, Mallikarjun B R, Shimon Vainer, Mark Boss, Yaoyao Liu
4Director makes video world models controllable by moving explicit 3D meshes through time while preserving realistic, consistent appearances.
EMPIRIC: Experiment-Driven Learning of Residual World Models for Robot Planning
Yichao Liang, Amber Li, Dat Nguyen, Emily Bunnapradist, Michelangelo Naim, Sreela Kodali, Matteo Merler, Bowen Li, Kiran Gopinathan, Yiyun Liu, Nikhil Pimpalkhare, Joshua B. Tenenbaum, Adrian Weller, Zenna Tavares, Tom Silver, Kevin Ellis
EMPIRIC lets robots discover missing physics through targeted experiments and use the resulting interpretable world models to plan better.
RoboCoach: World Models as Active Coaches for Compositional Robot Skills
Jiajun Liu, Yifan Chen, Yichao Liu, Jiayi Zhang, Ruoqu Chen, Shaoxuan Xie, Guocai Yao, Mengdi Xu, Sen Cui, Changshui Zhang
RoboCoach uses imagined robot failures to decide what demonstrations to request next, making long-horizon manipulation skills improve more efficiently.