ReViV: Reconstructing the Viewer and the View in 4D from Monocular Egocentric Video
AuthorsXiaozhong Lyu, Gen Li, Zhiyin Qian, Xucong Zhang, Marc Pollefeys, Siyu Tang
Resources
ReViV turns a single wearable-camera video into a coherent 4D model of both the person wearing it and the world they are viewing.
Key results
Unified multimodal corpus of unique tokens used to pretrain ReViV.
Total tokens processed during MGET training.
Seconds per clip for ReViV inference.
Pose error on the unseen Aria Digital Twin benchmark.
Hand reconstruction error in millimeters on HoloAssist.
Hand reconstruction error in millimeters on TACO.
What the paper found
ReViV, from ETH Zurich researchers including Marc Pollefeys of Microsoft Switzerland, presents a unified approach to reconstructing both the camera wearer and the surrounding scene in 4D from a single monocular egocentric RGB video. Its Masked Generative Egocentric Transformer, or MGET, converts RGB, depth, camera trajectory, gaze, hand motion, and full-body kinematics into modality-specific VQ-VAE tokens, then learns their joint distribution through masked multimodal prediction. A continuous Vision Transformer branch preserves fine visual detail lost during quantization, while floor-based least-squares alignment converts affine-invariant depth into a shared metric coordinate system. ReViV is pretrained on 7B unique tokens and optimized over 500B training tokens, covering datasets such as HoloAssist, HOT3D, ARCTIC, TACO, EgoExo4D, and Nymeria. On the unseen Aria Digital Twin benchmark, it reconstructs body motion using only RGB, reaching 88.6 PA-MPJPE and 0.442 FID, while running in 0.7 seconds per clip—100× faster than EgoAllo. For hand reconstruction, it achieves 10.5 millimeters PA-MPJPE on HoloAssist and 9.4 millimeters on TACO, outperforming HaMeR and Dyn-HaMR. ReViV also estimates gaze, camera motion, and depth competitively, although its depth accuracy trails EgoMono4D because discrete tokenization sacrifices fine geometric detail. The released system demonstrates that cross-modal generative pretraining can replace specialized SLAM, external hand trackers, and task-specific optimization in holistic egocentric perception.
Original abstract
Egocentric devices, such as wearable front-facing cameras, provide a unique perspective for capturing the continuous interaction between a human viewer and the surrounding environment. A holistic and efficient multimodal model capable of reconstructing this 4D representation is therefore highly desirable. However, existing approaches often rely on auxiliary inputs such as pre-computed camera trajectories, treat scene perception and human ego-motion modeling as separate problems despite their strong interdependency, and suffer from slow inference time. To address these limitations, we present ReViV, the first unified framework for holistic egocentric 4D reconstruction that extracts both viewer and view dynamics from a single monocular RGB video. We formulate the task as learning the full joint probability distribution over multimodal signals, including RGB video, camera trajectory, gaze direction, full-body motion, hand motion, and depth. Powered by a Masked Generative Egocentric Transformer, ReViV operates within a single feed-forward architecture to simultaneously reconstruct the temporally consistent 4D reconstruction across the viewer and the view with fast inference speed. Extensive experiments on diverse benchmarks, including HoloAssist, HOT3D, ARCTIC, Aria Digital Twin, and TACO, demonstrate that ReViV achieves state-of-the-art accuracy and efficiency across holistic ego-body, hand, and gaze reconstruction, camera tracking, while maintaining highly competitive egocentric depth estimation without relying on heavy task-specific priors. Code and models are fully open-sourced: https://reviv4d.github.io/.
Read the original paperMore in Embodied AI
Browse all 48 papers →GroundingPI: A Grounding Foundation Model towards Physical Intelligence with Visual Primitives
Qize Yu, Lianrui Fan, Boyu Chen, Jiaqi Liang, Xini Ding, Yue Chen, Zetian Song, Yuran Wang, Yi Zou, Kaixuan Wang, Tianxing Chen, Wenxuan Song, Bohan Zhou, Mingleyang Li, Siqiao Huang, Yuqi Ye, Caigao Jiang, Wei Wei, Ruihai Wu, Hang Zhang, Yixiao Ge, Shuchang Zhou, Shilong Liu, Xianming Liu, Ping Luo, Shiyu Huang
GroundingPI argues that fast, precise visual grounding should be the perceptual foundation for capable robots and autonomous vehicles.
MM-ABC: Towards Generalist Mobile Manipulation via Seeing, Coordinating and Imagining
Qiwei Liang, Guangyu Chen, Shaolong Zhu, Zikuan Xiao, Jinxuan Lu, Yifan Xie, Renjing Xu, Wenbo Ding, Tianxing Chen
MM-ABC is a generalist robot foundation model that helps mobile manipulators see their surroundings, coordinate arm and base motion, and imagine future actions for better performance.
Morphometric Imitation: From Morphology and Contact Aware Hand Retargeting to Sim-to-Real Visuomotor Policy
Tara Sadjadpour, Siming He, C. K. Wolfe, Haozhi Qi, Lea Wilken, S. Shankar Sastry, Claire Tomlin, Jitendra Malik
A three-stage system converts human hand demonstrations into robust, zero-shot real-robot dexterous manipulation policies across different hand morphologies.