NTH

EgoSteer: A Full-Stack System Towards Steerable Dexterous Manipulation from Egocentric Videos

AuthorsYifan Zhong, Zhang Chen, Tianrui Guan, Fanlian Zeng, Yuyao Ye, Tianjia He, Ka Nam Lui, Jiayi Li, Tingrui Zhang, Ruilin Yan, Xinhao Ji, Guangyu Zhao, Wenjie Lou, Jiayuan Zhang, Yuanpei Chen, Yaodong Yang

July 21, 2026 2 min read
Watch on YouTube
The one-line take

EgoSteer is a full-stack robot learning pipeline that turns massive egocentric human video into a steerable dexterous manipulation policy capable of generalizing to many real-world tasks.

Key results

9.6K
EgoSmith pre-training corpus

Hours of curated egocentric video across 12 datasets.

9×
EgoSmith throughput gain

Processing speed improvement over prior HaWoR-based processing.

75%
Multi-task success rate

Average success across 40 free-form dexterous manipulation tasks.

62.5%
DAgger improvement

Post-DAgger success rate, up from 22.5% on failure-prone tasks.

83%
Cake-unboxing adaptation

Few-shot success rate on AgiBot G1.

What the paper found

Researchers at Peking University’s Institute for AI and PKU-PsiBot Joint Lab introduce EgoSteer, a full-stack system for steerable dexterous manipulation from egocentric video. Its EgoSmith pipeline filters in-the-wild footage, reconstructs metric 4D hand motion with DPVO and Any4D, and generates five-level language annotations with Qwen3.5-VL-Plus, producing 9.6K hours across 12 datasets at 9× the throughput of prior HaWoR-based processing. The robot stack unifies PsiBot SynGlove-Air teleoperation, policy inference, and seamless human intervention through relative-motion mapping, collecting 187 hours across 193 dexterous tasks and enabling DAgger correction. EgoSteer combines a Qwen3-VL 2B backbone with a DiT flow-matching action expert, a training-only world-model expert that predicts future DINOv3 features, and training-time Real-Time Chunking to avoid execution pauses. On RealMan hardware, the system follows free-form instructions across 40 tasks with 75% average success, including 65% on compositional generalization and 62% on unseen tasks. Three DAgger iterations raise success on failure-prone tasks from 22.5% to 62.5%. Few-shot adaptation transfers to long-horizon box folding on RealMan at 75% success and cake unboxing on AgiBot G1 at 83%, while Diffusion Policy, IMLE, and a from-scratch model achieve 0%.

Original abstract

Steerability is a defining capability of generalist robot policies, yet remains largely absent in dexterous-hand systems for lack of large-scale, language-aligned, and action-accurate demonstration data. To address this bottleneck, we present a full-stack system that scales dexterous VLA pre-training from egocentric human videos and enables data-efficient real-robot post-training. It integrates EgoSmith, a data pipeline that curates in-the-wild egocentric videos into 9.6K hours of high-quality pre-training data with 9x higher throughput and better accuracy than prior SOTA; a unified robot stack for teleoperation and human-in-the-loop correction; and EgoSteer, a world-model-enhanced VLA trained on optimized infrastructure. Human-data pre-training equips EgoSteer with language-guided manipulation priors, which are grounded through robot post-training and improved by DAgger refinement. Empirically, EgoSteer robustly executes free-form instructions across 40+ diverse tasks, demonstrating failure recovery, dexterity, and generalization. The pre-trained model also few-shot adapts to complex long-horizon tasks, including box folding, on two embodiments with 75+% success. We open-source the system, data, and model at https://egosteer.github.io/.

Read the original paper

More in Robotics

Browse all 50 papers →
02Robotics

Rolling-WAM: World Action Models with Rolling Imagination

Yinghua Zhou, Junjie Ye, Yiqi Zhao, Hao Dong, Celina Shiyu Wang, Ruohai Ge, Tingyi Yang, Basile Van Hoorick, Gaurav Sukhatme, Vitor Guizilini, Yue Wang

Rolling-WAM keeps future robot actions partially imagined and refined over time, making world-model-based manipulation replan 4.5 times faster.

Read analysis
03Robotics

Training-free Behavior Cloning

Maximilian Adang, Timothy Chen, Lars Osterberg, Aiden Swann, Mac Schwager

A fast, training-free robot controller reuses and corrects demonstration trajectories to deliver traceable behavior at real-time speeds.

Read analysis