NTH

ReViV: Reconstructing the Viewer and the View in 4D from Monocular Egocentric Video

AuthorsXiaozhong Lyu, Gen Li, Zhiyin Qian, Xucong Zhang, Marc Pollefeys, Siyu Tang

July 28, 2026 2 min read
Watch on YouTube
The one-line take

ReViV turns a single wearable-camera video into a coherent 4D model of both the person wearing it and the world they are viewing.

Key results

7B
Pretraining dataset scale

Unified multimodal corpus of unique tokens used to pretrain ReViV.

500B
Training token budget

Total tokens processed during MGET training.

0.7
Inference time

Seconds per clip for ReViV inference.

88.6
Body PA-MPJPE

Pose error on the unseen Aria Digital Twin benchmark.

10.5
HoloAssist hand PA-MPJPE

Hand reconstruction error in millimeters on HoloAssist.

9.4
TACO hand PA-MPJPE

Hand reconstruction error in millimeters on TACO.

What the paper found

ReViV, from ETH Zurich researchers including Marc Pollefeys of Microsoft Switzerland, presents a unified approach to reconstructing both the camera wearer and the surrounding scene in 4D from a single monocular egocentric RGB video. Its Masked Generative Egocentric Transformer, or MGET, converts RGB, depth, camera trajectory, gaze, hand motion, and full-body kinematics into modality-specific VQ-VAE tokens, then learns their joint distribution through masked multimodal prediction. A continuous Vision Transformer branch preserves fine visual detail lost during quantization, while floor-based least-squares alignment converts affine-invariant depth into a shared metric coordinate system. ReViV is pretrained on 7B unique tokens and optimized over 500B training tokens, covering datasets such as HoloAssist, HOT3D, ARCTIC, TACO, EgoExo4D, and Nymeria. On the unseen Aria Digital Twin benchmark, it reconstructs body motion using only RGB, reaching 88.6 PA-MPJPE and 0.442 FID, while running in 0.7 seconds per clip—100× faster than EgoAllo. For hand reconstruction, it achieves 10.5 millimeters PA-MPJPE on HoloAssist and 9.4 millimeters on TACO, outperforming HaMeR and Dyn-HaMR. ReViV also estimates gaze, camera motion, and depth competitively, although its depth accuracy trails EgoMono4D because discrete tokenization sacrifices fine geometric detail. The released system demonstrates that cross-modal generative pretraining can replace specialized SLAM, external hand trackers, and task-specific optimization in holistic egocentric perception.

Original abstract

Egocentric devices, such as wearable front-facing cameras, provide a unique perspective for capturing the continuous interaction between a human viewer and the surrounding environment. A holistic and efficient multimodal model capable of reconstructing this 4D representation is therefore highly desirable. However, existing approaches often rely on auxiliary inputs such as pre-computed camera trajectories, treat scene perception and human ego-motion modeling as separate problems despite their strong interdependency, and suffer from slow inference time. To address these limitations, we present ReViV, the first unified framework for holistic egocentric 4D reconstruction that extracts both viewer and view dynamics from a single monocular RGB video. We formulate the task as learning the full joint probability distribution over multimodal signals, including RGB video, camera trajectory, gaze direction, full-body motion, hand motion, and depth. Powered by a Masked Generative Egocentric Transformer, ReViV operates within a single feed-forward architecture to simultaneously reconstruct the temporally consistent 4D reconstruction across the viewer and the view with fast inference speed. Extensive experiments on diverse benchmarks, including HoloAssist, HOT3D, ARCTIC, Aria Digital Twin, and TACO, demonstrate that ReViV achieves state-of-the-art accuracy and efficiency across holistic ego-body, hand, and gaze reconstruction, camera tracking, while maintaining highly competitive egocentric depth estimation without relying on heavy task-specific priors. Code and models are fully open-sourced: https://reviv4d.github.io/.

Read the original paper

More in Embodied AI

Browse all 48 papers →
01Embodied Ai

GroundingPI: A Grounding Foundation Model towards Physical Intelligence with Visual Primitives

Qize Yu, Lianrui Fan, Boyu Chen, Jiaqi Liang, Xini Ding, Yue Chen, Zetian Song, Yuran Wang, Yi Zou, Kaixuan Wang, Tianxing Chen, Wenxuan Song, Bohan Zhou, Mingleyang Li, Siqiao Huang, Yuqi Ye, Caigao Jiang, Wei Wei, Ruihai Wu, Hang Zhang, Yixiao Ge, Shuchang Zhou, Shilong Liu, Xianming Liu, Ping Luo, Shiyu Huang

GroundingPI argues that fast, precise visual grounding should be the perceptual foundation for capable robots and autonomous vehicles.

Read analysis