Sekai2: From World Exploration to Interactive World Modeling
AuthorsKang He, Wenshuo Peng, Zihui Gao, Jiaming Tan, Kaipeng Zhang, Yongtao Ge
Resources
Sekai2 is a large-scale video dataset designed to teach AI systems how worlds evolve across time, viewpoints, and camera trajectories.
Key results
Total number of real-world video clips in Sekai2.
Total hours of released footage.
Countries or regions represented in the dataset.
Hierarchical semantic intervals aligned to video timelines.
Share of total footage contributed by segments reaching 120 seconds.
Revisit-rich panoramic sequences with loops and non-linear trajectories.
What the paper found
Sekai2 is a real-world video dataset designed to address a core weakness in interactive world modeling: conventional video-text corpora such as WebVid and Panda-70M provide visual diversity but usually lack camera trajectories and temporally aligned semantics. The release contains 128,892 clips totaling 2,826 hours across 113 countries or regions, with every clip paired with camera poses and hierarchical descriptions that separate subject motion, environmental dynamics, static scene content, and camera behavior. Using OmniShotCut for shot-aware segmentation, ViPE with GeoCalib and DROID-SLAM for monocular camera reconstruction, and Kimi-K2.6 through an OpenAI-compatible service for annotation, the pipeline produces 649,597 temporally grounded segments. Long-horizon coverage is substantial: 51.4% of the footage comes from segments reaching the 120-second analysis limit. A distinctive subset of 982 panoramic sequences uses non-linear routes, loops, and revisits; geometrically verified loop closures refine 81.9% of the original panoramic trajectories, providing supervision for persistent scene memory and revisit consistency—an issue also studied by systems such as DeepMind’s Genie, WorldMem, and Infinite-World. In quality evaluations, Sekai2 achieves a LAION aesthetic score of 5.23 and mean optical-flow magnitude of 35.08, indicating visual quality comparable to established corpora but stronger local temporal variation. The dataset is therefore positioned as infrastructure for long-horizon video generation, camera-controllable synthesis, and pre-training world models that must distinguish viewpoint changes from genuine changes in the environment.
Original abstract
Video world models must capture how scenes evolve over time and across viewpoints. Training them for long-horizon generation and camera control therefore benefits from long videos paired with camera trajectories and temporally grounded semantics. Existing corpora rarely offer the three together: large-scale web video provides broad visual diversity but no trajectories or time-aligned text, while pose-annotated datasets are typically short-range or reconstruction-oriented. We introduce Sekai2, a multi-source real-world video dataset that carries the world-exploration footage of Sekai toward interactive world modeling. The release contains 128,892 clips totaling 2,826 hours from 10,428 source videos across 113 countries or regions, and is deliberately weighted toward sustained observation: under a common 120-second decomposition, 43,594 segments reach the full two minutes and account for 51.4% of all footage. Every clip includes a released camera trajectory and hierarchical annotations disentangling subject motion, environment dynamics, static scene content, and camera behavior, resulting in 649,597 temporally grounded segments. Crucially, we further introduce 982 panoramic sequences captured along non-linear trajectories with loops and revisits. These revisits provide repeated observations of the same locations across time and viewpoints, offering essential supervision for learning persistent scene representations, long-term spatial memory, and geometrically consistent world models. Corpus-scale analyses demonstrate complete pose-and-caption coverage, broad geographic and semantic diversity, varied camera trajectories, and highly non-redundant temporal descriptions. Together, these properties make Sekai2 a scalable resource for long-horizon video generation, camera-controllable synthesis, and interactive world-model pre-training.
Read the original paperMore in World Models
Browse all 41 papers →4Director: Controlling Video World Models with Rigid 3D Geometry
Wei Cao, Hao Zhang, Vikram Voleti, Yuqun Wu, Mallikarjun B R, Shimon Vainer, Mark Boss, Yaoyao Liu
4Director makes video world models controllable by moving explicit 3D meshes through time while preserving realistic, consistent appearances.
EMPIRIC: Experiment-Driven Learning of Residual World Models for Robot Planning
Yichao Liang, Amber Li, Dat Nguyen, Emily Bunnapradist, Michelangelo Naim, Sreela Kodali, Matteo Merler, Bowen Li, Kiran Gopinathan, Yiyun Liu, Nikhil Pimpalkhare, Joshua B. Tenenbaum, Adrian Weller, Zenna Tavares, Tom Silver, Kevin Ellis
EMPIRIC lets robots discover missing physics through targeted experiments and use the resulting interpretable world models to plan better.
RoboCoach: World Models as Active Coaches for Compositional Robot Skills
Jiajun Liu, Yifan Chen, Yichao Liu, Jiayi Zhang, Ruoqu Chen, Shaoxuan Xie, Guocai Yao, Mengdi Xu, Sen Cui, Changshui Zhang
RoboCoach uses imagined robot failures to decide what demonstrations to request next, making long-horizon manipulation skills improve more efficiently.