NTH

Sekai2: From World Exploration to Interactive World Modeling

AuthorsKang He, Wenshuo Peng, Zihui Gao, Jiaming Tan, Kaipeng Zhang, Yongtao Ge

August 12, 2026 3 min read
Watch on YouTube
The one-line take

Sekai2 is a large-scale video dataset designed to teach AI systems how worlds evolve across time, viewpoints, and camera trajectories.

Key results

128,892
Released clips

Total number of real-world video clips in Sekai2.

2,826
Video duration

Total hours of released footage.

113
Geographic coverage

Countries or regions represented in the dataset.

649,597
Temporally grounded segments

Hierarchical semantic intervals aligned to video timelines.

51.4%
Full-length 120-second footage share

Share of total footage contributed by segments reaching 120 seconds.

982
Panoramic sequences

Revisit-rich panoramic sequences with loops and non-linear trajectories.

What the paper found

Sekai2 is a real-world video dataset designed to address a core weakness in interactive world modeling: conventional video-text corpora such as WebVid and Panda-70M provide visual diversity but usually lack camera trajectories and temporally aligned semantics. The release contains 128,892 clips totaling 2,826 hours across 113 countries or regions, with every clip paired with camera poses and hierarchical descriptions that separate subject motion, environmental dynamics, static scene content, and camera behavior. Using OmniShotCut for shot-aware segmentation, ViPE with GeoCalib and DROID-SLAM for monocular camera reconstruction, and Kimi-K2.6 through an OpenAI-compatible service for annotation, the pipeline produces 649,597 temporally grounded segments. Long-horizon coverage is substantial: 51.4% of the footage comes from segments reaching the 120-second analysis limit. A distinctive subset of 982 panoramic sequences uses non-linear routes, loops, and revisits; geometrically verified loop closures refine 81.9% of the original panoramic trajectories, providing supervision for persistent scene memory and revisit consistency—an issue also studied by systems such as DeepMind’s Genie, WorldMem, and Infinite-World. In quality evaluations, Sekai2 achieves a LAION aesthetic score of 5.23 and mean optical-flow magnitude of 35.08, indicating visual quality comparable to established corpora but stronger local temporal variation. The dataset is therefore positioned as infrastructure for long-horizon video generation, camera-controllable synthesis, and pre-training world models that must distinguish viewpoint changes from genuine changes in the environment.

Original abstract

Video world models must capture how scenes evolve over time and across viewpoints. Training them for long-horizon generation and camera control therefore benefits from long videos paired with camera trajectories and temporally grounded semantics. Existing corpora rarely offer the three together: large-scale web video provides broad visual diversity but no trajectories or time-aligned text, while pose-annotated datasets are typically short-range or reconstruction-oriented. We introduce Sekai2, a multi-source real-world video dataset that carries the world-exploration footage of Sekai toward interactive world modeling. The release contains 128,892 clips totaling 2,826 hours from 10,428 source videos across 113 countries or regions, and is deliberately weighted toward sustained observation: under a common 120-second decomposition, 43,594 segments reach the full two minutes and account for 51.4% of all footage. Every clip includes a released camera trajectory and hierarchical annotations disentangling subject motion, environment dynamics, static scene content, and camera behavior, resulting in 649,597 temporally grounded segments. Crucially, we further introduce 982 panoramic sequences captured along non-linear trajectories with loops and revisits. These revisits provide repeated observations of the same locations across time and viewpoints, offering essential supervision for learning persistent scene representations, long-term spatial memory, and geometrically consistent world models. Corpus-scale analyses demonstrate complete pose-and-caption coverage, broad geographic and semantic diversity, varied camera trajectories, and highly non-redundant temporal descriptions. Together, these properties make Sekai2 a scalable resource for long-horizon video generation, camera-controllable synthesis, and interactive world-model pre-training.

Read the original paper

More in World Models

Browse all 41 papers →
01World Model

4Director: Controlling Video World Models with Rigid 3D Geometry

Wei Cao, Hao Zhang, Vikram Voleti, Yuqun Wu, Mallikarjun B R, Shimon Vainer, Mark Boss, Yaoyao Liu

4Director makes video world models controllable by moving explicit 3D meshes through time while preserving realistic, consistent appearances.

Read analysis
02World Model

EMPIRIC: Experiment-Driven Learning of Residual World Models for Robot Planning

Yichao Liang, Amber Li, Dat Nguyen, Emily Bunnapradist, Michelangelo Naim, Sreela Kodali, Matteo Merler, Bowen Li, Kiran Gopinathan, Yiyun Liu, Nikhil Pimpalkhare, Joshua B. Tenenbaum, Adrian Weller, Zenna Tavares, Tom Silver, Kevin Ellis

EMPIRIC lets robots discover missing physics through targeted experiments and use the resulting interpretable world models to plan better.

Read analysis
03World Model

RoboCoach: World Models as Active Coaches for Compositional Robot Skills

Jiajun Liu, Yifan Chen, Yichao Liu, Jiayi Zhang, Ruoqu Chen, Shaoxuan Xie, Guocai Yao, Mengdi Xu, Sen Cui, Changshui Zhang

RoboCoach uses imagined robot failures to decide what demonstrations to request next, making long-horizon manipulation skills improve more efficiently.

Read analysis