Kairos: A Native World Model Stack for Physical AI
AuthorsKairos Team, Fei Wang, Shan You, Qiming Zhang, Tao Huang, Zuoyi Fu, Zhisheng Zheng, Yunlong Xi, Feng Lv, Xiaoming Wu, Zeyu Liu, Cong Wan, Pu Li, Ruiqing Yang, Xiaoou Li, Wei Wang, Kangkang Zhu, Yuwei Zhang, Shi Fu, Zheng Zhang, Xiaoning Wu, Xuzeng Fan, Dacheng Tao, Xiaogang Wang
Resources
Kairos proposes a full-stack world model for physical AI that learns from mixed embodied data, remembers over long horizons, and runs efficiently on real hardware.
Key results
Kairos-4B is the compact world-action model evaluated across the paper's benchmarks.
Kairos achieves the highest reported total score on the robotics subset.
Average success rate across clean and randomized bimanual manipulation settings.
Average score when future video and action tokens are jointly denoised.
Kairos distills a 480P embodied world model into a 4-step generator.
PFlops required for the reported 720P, five-second inference comparison.
What the paper found
Kairos proposes a regret-aware world-action stack for Physical AI, arguing that robots need compact control-sufficient states—not pixel-perfect simulations—that preserve object state, contact conditions, task progress, action consequences, failure risk, and deployment uncertainty. Positioned alongside NVIDIA Cosmos, Meta’s V-JEPA 2, and OpenAI’s Sora, Kairos unifies understanding, generation, and prediction through a shared latent state, using Video DiT and Action DiT to jointly model future visual and action tokens. Its Cross-Embodiment Data Curriculum progresses from passive internet video to human behavior and robot interaction, while Hybrid Linear Temporal Attention combines Sliding-Window Attention for local contact dynamics, Dilated Sliding-Window Attention for mid-range events, and Gated Linear Attention with GatedDeltaNet for persistent global memory. Flow Matching, DPO-style regret alignment, timestep distillation, quantization, and hardware-aware kernels target deployment efficiency. The compact 4B model scores 9.30 on WorldModelBench-Robot, reaches 96.1 average success on RoboTwin 2.0, and achieves 90.8 on LIBERO-Plus when future video and action prediction are jointly denoised; its 4-step distilled generator reduces sampling cost, while Kairos-4B requires 2.3 PFlops for 720P five-second inference. At 15 seconds, it records a 79.9 PAI-Bench overall score. These are proxy results rather than proof of lower real-world regret: imagined-to-real rollout correlation, counterfactual action validation, failure prediction, safety filtering, recovery learning, and real-robot policy improvement remain future evaluations.
Original abstract
World models are transitioning from passive visual generators to foundational, operational infrastructure for Physical AI: they must natively acquire world knowledge from heterogeneous experience, maintain persistent states over long horizons, and execute efficiently within real deployment constraints. We introduce Kairos, a native world model stack designed around these requirements. (1) Kairos learns the world by pioneering a Native Pre-training Paradigm governed by a Cross-Embodiment Data Curriculum, which organizes open-world videos, human behavioral data, and robot interactions into a progressive developmental pathway. (2) Kairos maintains the world by unified world understanding, generation, and prediction within a Native Unified Architecture equipped with Hybrid Linear Temporal Attention, where sliding-window attention captures local dynamics, dilated sliding windows capture mid-range dependencies, and gated linear attention maintains persistent global memory. We establish formal theoretical bounds demonstrating that this temporal factorization strictly limits error accumulation, mathematically guaranteeing state propagation across extended horizons. (3) Kairos runs the world by incorporating a Deployment-Aware System Co-Design to support low-latency rollout generation on server and consumer-grade hardware for real-world observation-action-feedback loops. Experiments on embodied world-model, long-horizon, and action-policy benchmarks show that Kairos achieves top level performance while offering a strong efficiency-capability trade-off. Together, these results position Kairos as a cohesive operational foundation for future self-evolving physical intelligence.
Read the original paperMore in World Models
Browse all 41 papers →4Director: Controlling Video World Models with Rigid 3D Geometry
Wei Cao, Hao Zhang, Vikram Voleti, Yuqun Wu, Mallikarjun B R, Shimon Vainer, Mark Boss, Yaoyao Liu
4Director makes video world models controllable by moving explicit 3D meshes through time while preserving realistic, consistent appearances.
EMPIRIC: Experiment-Driven Learning of Residual World Models for Robot Planning
Yichao Liang, Amber Li, Dat Nguyen, Emily Bunnapradist, Michelangelo Naim, Sreela Kodali, Matteo Merler, Bowen Li, Kiran Gopinathan, Yiyun Liu, Nikhil Pimpalkhare, Joshua B. Tenenbaum, Adrian Weller, Zenna Tavares, Tom Silver, Kevin Ellis
EMPIRIC lets robots discover missing physics through targeted experiments and use the resulting interpretable world models to plan better.
RoboCoach: World Models as Active Coaches for Compositional Robot Skills
Jiajun Liu, Yifan Chen, Yichao Liu, Jiayi Zhang, Ruoqu Chen, Shaoxuan Xie, Guocai Yao, Mengdi Xu, Sen Cui, Changshui Zhang
RoboCoach uses imagined robot failures to decide what demonstrations to request next, making long-horizon manipulation skills improve more efficiently.