NTH

HuRo: Robotizing Human Videos for Scalable VLA Pretraining

AuthorsJinho Jeong, Se June Joo, Jaehyun Kang, Dongyun Kim, Yena Kim, Hanjung Kim, Seon Joo Kim

AffiliationsRLWRLD · Yonsei University

September 21, 2026 2 min read
Watch on YouTube
The one-line take

HuRo turns massive human-video collections into robot-training data, substantially improving the robustness and performance of vision-language-action policies.

Key results

630K
HuRo episodes

Robotized episodes generated from five human-video sources.

142M
Processed frames

Total frames in the HuRo pretraining dataset.

80.3%
Overall completion

Four-task completion with the full HuRo pretraining dataset, up from 51.5% without HuRo.

72.2%
OOD completion

Completion under spatial and visual distribution shifts with full HuRo pretraining.

50.0%
End-to-end VLA OOD completion

Diverse Pick-and-Place performance using visual-plus-action pretraining.

What the paper found

HuRo proposes a pipeline for turning heterogeneous egocentric human videos into robot-aligned supervision for vision-language-action, or VLA, pretraining. It estimates camera geometry and hand pose, captions manipulation chunks with Qwen3.5, retargets motion into robot trajectories with PyRoKi, removes human arms using SAM2 and ProPainter, and overlays the target robot using NVIDIA Isaac Sim. The resulting HuRo dataset combines Ego4D, EPIC-Kitchens, EgoDex, EgoVerse, and Ego10K into 630K robotized episodes and 142M processed frames. Policies built on the GR00T-N1.6-3B architecture are pretrained end-to-end on robotized observations and retargeted actions, then fine-tuned on four real-world ALLEX manipulation tasks. Scaling pretraining from no HuRo data to the full dataset raises overall completion from 51.5% to 80.3%, while out-of-distribution completion under spatial and visual shifts rises from 34.9% to 72.2%. Visual robotization is essential for generalization: a full no-overlay variant reaches 55.7% OOD completion, versus 72.2% with robot overlays. Action supervision also matters: end-to-end visual-plus-action pretraining reaches 61.1% in-distribution and 50.0% OOD completion on Diverse Pick-and-Place, outperforming visual-only transfer. Compared with video-generation pretraining using I2V plus an inverse-dynamics model, HuRo at 0.7M frames already exceeds the baseline at 7.0M frames, indicating more favorable scaling.

Original abstract

Human video datasets have emerged as a compelling alternative to expensive real-robot data, offering rich diversity at scale. To bridge the human-to-robot embodiment gap, existing approaches either robotize videos in task-matched settings or address observation and action alignment separately at scale. In this work, we systematically examine whether robotized human videos can provide effective and scalable supervision for pretraining vision-language-action (VLA) policies. To this end, we develop a robotization pipeline that converts heterogeneous human videos into robot-aligned observations and action trajectories while inferring missing intermediate signals across annotation levels. Using this pipeline, we construct the HuRo dataset, comprising about 630K robotized episodes and 142M processed frames from five human-video sources. Across four real-world manipulation tasks, increasing robotized pretraining scale improves overall completion from 51.5% to 80.3% and OOD completion under spatial and visual shifts from 34.9% to 72.2%. Ablations further show that visual robotization improves OOD robustness and that end-to-end pretraining with retargeted actions outperforms visual-only transfer. Code and data are released on our website: https://3587jjh.github.io/HuRo.

Read the original paper

More in Embodied AI

Browse all 48 papers →
01Embodied Ai

GroundingPI: A Grounding Foundation Model towards Physical Intelligence with Visual Primitives

Qize Yu, Lianrui Fan, Boyu Chen, Jiaqi Liang, Xini Ding, Yue Chen, Zetian Song, Yuran Wang, Yi Zou, Kaixuan Wang, Tianxing Chen, Wenxuan Song, Bohan Zhou, Mingleyang Li, Siqiao Huang, Yuqi Ye, Caigao Jiang, Wei Wei, Ruihai Wu, Hang Zhang, Yixiao Ge, Shuchang Zhou, Shilong Liu, Xianming Liu, Ping Luo, Shiyu Huang

GroundingPI argues that fast, precise visual grounding should be the perceptual foundation for capable robots and autonomous vehicles.

Read analysis