NTH

Robots Acquire Manipulation Skills in Seconds from a Single Human Video

AuthorsGuangyan Chen, Meiling Wang, Te Cui, Zichen Zhou, Qi Shao, Shalfun Li, Hang Su, Roy Gan, Hao Wang, Mengyin Fu, Yi Yang, Yufeng Yue

July 28, 2026 3 min read
Watch on YouTube
The one-line take

HOST lets robots learn new manipulation skills from a single human video in under a minute while preserving what they already know.

Key results

62%
Novel-task success rate

Average success from a single human video on novel manipulation tasks.

50
Novel manipulation tasks

Number of unseen tasks evaluated on the bimanual robot platform.

29
Inference-time skill acquisition

Average seconds required to make each new skill executable.

507
Acquisition speedup

Times faster than the fastest supervised fine-tuning baseline.

193462
Stage 1 pretraining trajectories

Robot trajectories used for same-embodiment pretraining across 229 tasks.

What the paper found

Researchers at Beijing Institute of Technology, X SQUARE ROBOT, and Tsinghua University introduce HOST, or Human-to-robot One-Shot Skill Acquisition, a framework that moves robot teaching from costly training-time fine-tuning to inference time. Given a single human video, HOST first uses self-supervised temporal alignment—combining Smooth Dynamic Time Warping and temporal cycle consistency—to locate the robot’s current stage in the demonstrated procedure. A causal self-grounded cascade then predicts the robot’s future observations and derives actions from those predictions, adapting behavior across differences in embodiment, viewpoint, objects, and scene. Its dual-expert Mixture-of-Transformers uses a Wan2.2 video diffusion backbone, Qwen3-VL-Embedding-8B for alignment, and flow matching for localization, visual prediction, and action generation. On 50 novel manipulation tasks using an ARX R5 bimanual platform, HOST reaches 62% average success from one video, exceeding the strongest zero-shot baseline by 45%. Each skill becomes executable in 29 seconds on average, versus roughly four hours for supervised fine-tuning. Against a baseline trained on 50 robot demonstrations per task, HOST uses 50 times fewer demonstrations and acquires skills 507 times faster, while preserving previously mastered behaviors because its policy parameters remain frozen. The training pipeline uses 193462 robot trajectories across 229 tasks for same-embodiment pretraining and 5847 human–robot pairs for adaptation, after which acquired videos can be stored and retrieved for recurring tasks.

Original abstract

The ability to acquire skills rapidly and effortlessly while retaining those already mastered is essential for robots. However, current methods still rely on a cumbersome training-time loop that is costly and slow, while eroding skills already mastered. In this paper, we introduce HOST (Human-to-robot One-Shot Skill AcquisiTion), a framework that enables a robot to acquire skills in seconds from a single human video while retaining previously mastered skills. HOST resolves skill acquisition through a cascade of self-grounded prediction. It first estimates the robot's progress within the demonstrated task, then translates the upcoming progression into the robot's own future observations, and finally derives actions from these predicted observations. This cascade is trained on targets coupled to the video demonstration, obtained by mapping the robot trajectory and the video demonstration onto a shared task progress manifold, then redefining each target to align with the future progression of the video. HOST thereby enables the robot to actively follow the demonstrated procedure and adapt it to the robot's embodiment. HOST acquires novel skills at inference time from a single human video in an average of 29 seconds and achieves a 62% average success rate. It exceeds the zero-shot baseline by 45% while retaining previously mastered skills. HOST even exceeds the baseline fine-tuned on 50 robot demonstrations per task while requiring 50 times fewer demonstrations and acquiring each skill 507 times faster. Additional information about HOST is available on the project website.

Read the original paper

More in Embodied AI

Browse all 48 papers →
01Embodied Ai

GroundingPI: A Grounding Foundation Model towards Physical Intelligence with Visual Primitives

Qize Yu, Lianrui Fan, Boyu Chen, Jiaqi Liang, Xini Ding, Yue Chen, Zetian Song, Yuran Wang, Yi Zou, Kaixuan Wang, Tianxing Chen, Wenxuan Song, Bohan Zhou, Mingleyang Li, Siqiao Huang, Yuqi Ye, Caigao Jiang, Wei Wei, Ruihai Wu, Hang Zhang, Yixiao Ge, Shuchang Zhou, Shilong Liu, Xianming Liu, Ping Luo, Shiyu Huang

GroundingPI argues that fast, precise visual grounding should be the perceptual foundation for capable robots and autonomous vehicles.

Read analysis