ActiveMimic: Egocentric Video Pretraining with Active Perception
AuthorsXingyao Lin, Guojin Zhong, Tianyi Lu, Ziyi Ye, Yichen Zhu, Zuxuan Wu, Yu-Gang Jiang
Resources
ActiveMimic teaches robots from human first-person videos by treating camera motion as useful action, helping bridge the gap between egocentric video and robot pretraining.
Key results
End-to-end success rate on the Restocking task
End-to-end success rate on the Reaching task
End-to-end success rate on the Finding task
End-to-end success rate on the Pouring task
Filtered Ego4D episodes used for egocentric pretraining
Strict-tier recovery rate for predicted head pose on HOT3D
What the paper found
ActiveMimic, from Fudan University, reframes egocentric video pretraining for robot manipulation around active perception: instead of treating head-mounted camera motion as noise, it recovers synchronized camera and bimanual wrist trajectories from a single RGB egocentric stream, aligns them in a common reference frame, and encodes them as a unified 27D action space. The method uses off-the-shelf pose models, including SAM-3D-Body, VGGT, and UniDepth, and trains a 3B visual-language prefix plus a 0.6B action expert with flow matching on filtered Ego4D manipulation clips. On a humanoid AGI-BOT G1 across four real-world tasks, ActiveMimic reaches 90.1% on Restocking, 88.9% on Reaching, 91.7% on Finding, and 93.3% on Pouring, outperforming human-video baselines such as MotoVLA and matching or beating π0, a robot-data-pretrained model. Its pretraining corpus contains 2,561 Ego4D episodes, roughly 10 hours at 10 fps. On HOT3D, the recovered labels achieve 78.82% head recovery under the strict tolerance tier, with 65.93% and 61.72% for the left and right wrist. Analysis shows the active-perception capability comes from egocentric pretraining, not robot fine-tuning: on Restocking placement, ActiveMimic scores 24 out of 27, but only 1 out of 27 when the head camera is removed, and camera-motion supervision yields consistently higher head-view/full-view activation overlap in early-to-mid layers.
Original abstract
Egocentric human video offers a scalable alternative to robot data for pretraining, yet models pretrained on such video consistently underperform those pretrained on robot data. We attribute this gap to a missing signal, the active perception behavior in egocentric videos, where humans continuously reposition their viewpoint during manipulation, inducing camera motion that standard pipelines treat as noise. To address this, we present ActiveMimic, a pretraining framework that recovers synchronized camera and wrist trajectories from a single body-worn RGB camera, models camera motion as a viewpoint action, and jointly learns active perception and manipulation from in-the-wild egocentric human video before adapting to a target robot. Empirically, real-world experiments across tasks with diverse active perception demands show that ActiveMimic consistently surpasses baselines pretrained on human video and matches state-of-the-art models pretrained on robot data. Further analysis provides evidence that active perception capability originates from egocentric human video pretraining rather than robot-specific fine-tuning, confirming active perception as the key to unlocking egocentric human video for robot pretraining.
Read the original paperMore in Robotics
Browse all 50 papers →JAMB: Joint Action-Motion Diffusion for Bimanual Manipulation
Chuyang Xiao, Peilin Meng, David Held
JAMB helps two robot arms coordinate by jointly imagining their future movements and the changing 3D scene before acting.
Rolling-WAM: World Action Models with Rolling Imagination
Yinghua Zhou, Junjie Ye, Yiqi Zhao, Hao Dong, Celina Shiyu Wang, Ruohai Ge, Tingyi Yang, Basile Van Hoorick, Gaurav Sukhatme, Vitor Guizilini, Yue Wang
Rolling-WAM keeps future robot actions partially imagined and refined over time, making world-model-based manipulation replan 4.5 times faster.
Training-free Behavior Cloning
Maximilian Adang, Timothy Chen, Lars Osterberg, Aiden Swann, Mac Schwager
A fast, training-free robot controller reuses and corrects demonstration trajectories to deliver traceable behavior at real-time speeds.