EgoSteer: A Full-Stack System Towards Steerable Dexterous Manipulation from Egocentric Videos
AuthorsYifan Zhong, Zhang Chen, Tianrui Guan, Fanlian Zeng, Yuyao Ye, Tianjia He, Ka Nam Lui, Jiayi Li, Tingrui Zhang, Ruilin Yan, Xinhao Ji, Guangyu Zhao, Wenjie Lou, Jiayuan Zhang, Yuanpei Chen, Yaodong Yang
Resources
EgoSteer is a full-stack robot learning pipeline that turns massive egocentric human video into a steerable dexterous manipulation policy capable of generalizing to many real-world tasks.
Key results
Hours of curated egocentric video across 12 datasets.
Processing speed improvement over prior HaWoR-based processing.
Average success across 40 free-form dexterous manipulation tasks.
Post-DAgger success rate, up from 22.5% on failure-prone tasks.
Few-shot success rate on AgiBot G1.
What the paper found
Researchers at Peking University’s Institute for AI and PKU-PsiBot Joint Lab introduce EgoSteer, a full-stack system for steerable dexterous manipulation from egocentric video. Its EgoSmith pipeline filters in-the-wild footage, reconstructs metric 4D hand motion with DPVO and Any4D, and generates five-level language annotations with Qwen3.5-VL-Plus, producing 9.6K hours across 12 datasets at 9× the throughput of prior HaWoR-based processing. The robot stack unifies PsiBot SynGlove-Air teleoperation, policy inference, and seamless human intervention through relative-motion mapping, collecting 187 hours across 193 dexterous tasks and enabling DAgger correction. EgoSteer combines a Qwen3-VL 2B backbone with a DiT flow-matching action expert, a training-only world-model expert that predicts future DINOv3 features, and training-time Real-Time Chunking to avoid execution pauses. On RealMan hardware, the system follows free-form instructions across 40 tasks with 75% average success, including 65% on compositional generalization and 62% on unseen tasks. Three DAgger iterations raise success on failure-prone tasks from 22.5% to 62.5%. Few-shot adaptation transfers to long-horizon box folding on RealMan at 75% success and cake unboxing on AgiBot G1 at 83%, while Diffusion Policy, IMLE, and a from-scratch model achieve 0%.
Original abstract
Steerability is a defining capability of generalist robot policies, yet remains largely absent in dexterous-hand systems for lack of large-scale, language-aligned, and action-accurate demonstration data. To address this bottleneck, we present a full-stack system that scales dexterous VLA pre-training from egocentric human videos and enables data-efficient real-robot post-training. It integrates EgoSmith, a data pipeline that curates in-the-wild egocentric videos into 9.6K hours of high-quality pre-training data with 9x higher throughput and better accuracy than prior SOTA; a unified robot stack for teleoperation and human-in-the-loop correction; and EgoSteer, a world-model-enhanced VLA trained on optimized infrastructure. Human-data pre-training equips EgoSteer with language-guided manipulation priors, which are grounded through robot post-training and improved by DAgger refinement. Empirically, EgoSteer robustly executes free-form instructions across 40+ diverse tasks, demonstrating failure recovery, dexterity, and generalization. The pre-trained model also few-shot adapts to complex long-horizon tasks, including box folding, on two embodiments with 75+% success. We open-source the system, data, and model at https://egosteer.github.io/.
Read the original paperMore in Robotics
Browse all 50 papers →JAMB: Joint Action-Motion Diffusion for Bimanual Manipulation
Chuyang Xiao, Peilin Meng, David Held
JAMB helps two robot arms coordinate by jointly imagining their future movements and the changing 3D scene before acting.
Rolling-WAM: World Action Models with Rolling Imagination
Yinghua Zhou, Junjie Ye, Yiqi Zhao, Hao Dong, Celina Shiyu Wang, Ruohai Ge, Tingyi Yang, Basile Van Hoorick, Gaurav Sukhatme, Vitor Guizilini, Yue Wang
Rolling-WAM keeps future robot actions partially imagined and refined over time, making world-model-based manipulation replan 4.5 times faster.
Training-free Behavior Cloning
Maximilian Adang, Timothy Chen, Lars Osterberg, Aiden Swann, Mac Schwager
A fast, training-free robot controller reuses and corrects demonstration trajectories to deliver traceable behavior at real-time speeds.