NTH

SolarWM: Open Data and Scalable Training for Long-Horizon Video World Models

AuthorsJunchao Huang, Guian Fang, Shengju Qian, Xianghao Kong, Zhuoran Zhao, Wei Huang, Yihua Du, Zixin Zhang, Justin Cui, Yuchao Gu, Yukang Chen, Xinting Hu, Tianyu He, Shaoshuai Shi, Zhuotao Tian, Xin Wang, Mike Zheng Shou, Li Jiang

September 7, 2026 2 min read
Watch on YouTube
The one-line take

SolarWM is an open, scalable toolkit for training video world models that can simulate interactive environments over minutes or even hours.

Key results

1.43M
Canonical clips

Unified clips processed by the SolarWM open data engine.

10
Source datasets

Real-world, synthetic, and game datasets contributing to the corpus.

25.85
Corpus storage

Physical storage size in TB for the complete corpus.

33B
Largest model

Parameter count of the SolarWM-minimax-h3-33B model.

5
Training sequence length

Seconds used for training sequences.

60
Longest demonstrated rollout

Minutes reached by uninterrupted autoregressive generation.

What the paper found

SolarWM presents an open, reconfigurable foundation for interactive video world models, addressing the field’s central mismatch between heterogeneous video data and incompatible generation backbones. Its data engine converts 1.43M canonical clips from 10 datasets, spanning real, synthetic, and game environments, into frame-aligned records containing video, metric camera geometry, dense captions, quality scores, filtering decisions, and provenance, with the corpus occupying 25.85 TB. The framework supports four backbone-native models built on Wan2.2, LTX-2.5, and MiniMax-H3, including a largest 33B-parameter variant, while related open systems include Google’s Genie 3, Microsoft’s MineWorld, and NVIDIA’s SANA-WM. Training uses three stages: bidirectional camera-conditioned adaptation with fused-PRoPE, teacher-forced AnyFlow autoregressive initialization, and distribution-matching distillation, eliminating specialized ODE or consistency-distillation initialization. The causal models use four sampling steps at 16 fps and are trained only on 5s sequences, yet produce real-time, uninterrupted rollouts reaching a demonstrated 60-minute horizon without attention sinks, long-sequence fine-tuning, reference-frame resets, or clip splicing. The main contribution is not a single architecture, but a reproducible stack that separates expensive source preprocessing from configurable data mixtures while preserving each backbone’s native representation and objective.

Original abstract

We introduce SolarWM, a fully open foundation for building interactive video world models from data preparation through long-horizon inference. Training across heterogeneous data sources and video backbones is challenging: datasets differ in temporal scale, camera geometry, visual quality, motion, and captioning styles, while video generators use distinct representations and architectures. Naive data mixing and model-specific implementations therefore produce inconsistent supervision and make results difficult to reproduce and compare. SolarWM addresses this coupling with a reconfigurable multi-source data engine and a backbone-native adaptation framework. The engine converts 1.43 million canonical clips from 10 datasets into a unified, frame-aligned contract covering visual observations, metric camera geometry, captions, quality metadata, selection decisions, and provenance, while decoupling source processing from mixture construction. Under shared camera-conditioning, training, and inference interfaces, we instantiate four 5B--33B models based on Wan2.2, LTX-2.5, and MiniMax-H3 while preserving their native representations and objectives. A unified three-stage recipe combines bidirectional adaptation, teacher-forced autoregressive initialization, and distribution matching distillation. The resulting causal models enable real-time interaction over rollouts ranging from minutes to hours after being trained on only 5s sequences. By releasing the resulting data, pipeline, recipes, weights, and framework, SolarWM provides a reproducible and extensible foundation for interactive world-model research.

Read the original paper

More in World Models

Browse all 41 papers →
01World Model

4Director: Controlling Video World Models with Rigid 3D Geometry

Wei Cao, Hao Zhang, Vikram Voleti, Yuqun Wu, Mallikarjun B R, Shimon Vainer, Mark Boss, Yaoyao Liu

4Director makes video world models controllable by moving explicit 3D meshes through time while preserving realistic, consistent appearances.

Read analysis
02World Model

EMPIRIC: Experiment-Driven Learning of Residual World Models for Robot Planning

Yichao Liang, Amber Li, Dat Nguyen, Emily Bunnapradist, Michelangelo Naim, Sreela Kodali, Matteo Merler, Bowen Li, Kiran Gopinathan, Yiyun Liu, Nikhil Pimpalkhare, Joshua B. Tenenbaum, Adrian Weller, Zenna Tavares, Tom Silver, Kevin Ellis

EMPIRIC lets robots discover missing physics through targeted experiments and use the resulting interpretable world models to plan better.

Read analysis
03World Model

RoboCoach: World Models as Active Coaches for Compositional Robot Skills

Jiajun Liu, Yifan Chen, Yichao Liu, Jiayi Zhang, Ruoqu Chen, Shaoxuan Xie, Guocai Yao, Mengdi Xu, Sen Cui, Changshui Zhang

RoboCoach uses imagined robot failures to decide what demonstrations to request next, making long-horizon manipulation skills improve more efficiently.

Read analysis