HandEdit: A Unified Benchmark for Egocentric Human-to-Robot Dexterous Hand Image Editing
AuthorsZhenjie Yang, Xingyu Jiao, Guopeng Zhong, Shuzhe Yang, Shi Che, Chao Wu, Chenyu Jiang, Dongjie Zhang, Yideng Zhang, Zheng Zhang, Muyun Jiang, Haisheng Su, Shuang Jin, Donghang Zhang, Chao Yang, Li Chen, Hongyang Li, Zuxuan Wu, Yu-Gang Jiang, Xiaosong Jia, Junchi Yan
Resources
HandEdit turns abundant human hand videos into training data for dexterous robots by benchmarking how well image-editing models can adapt them to different robotic hands.
Key results
Image-level human-to-robot editing instances in HandEdit
URDF configurations covering 13 hand-only and 13 hand-arm embodiments
Commercial and open-source image-editing models benchmarked
Best Hand-only structural fidelity score
Highest Hand-only GPT-4o VLM judgment score
What the paper found
HandEdit introduces a benchmark for converting egocentric human hand and hand-arm manipulation images into URDF-specified dexterous robot embodiments. Built from five source datasets—EgoDex, ARCTIC, OakInk2, HOI4D, and HO-Cap—the release contains over 200M image-level editing instances spanning 26 target configurations: 13 hand-only and 13 hand-arm embodiments. Its curation pipeline uses SAM3 for segmentation, ProPainter for background restoration, MANO or 3D hand pose retargeting with inverse kinematics, URDF rendering, compositing, and Harmonizer-based appearance correction. The benchmark defines Hand-only and Hand-Arm tracks, each with 1K test images, and evaluates generic similarity, GPT-4o VLM judgments, and embodiment-aware measures for human-hand removal, structural and identity fidelity, and interaction preservation. Across 11 commercial and open-source editors, OpenAI’s GPT-Image-2 is the strongest overall baseline: it reaches 0.780 structural fidelity and 0.703 interaction consistency on Hand-only, while OpenAI’s GPT-Image-1.5 achieves the highest Hand-only VLM score of 0.765. Google’s Nano-Banana-2 and ByteDance’s Seedream-4.5 are also evaluated, but the results show that visually plausible editing does not guarantee correct robot morphology or preserved hand-object contact. HandEdit’s central contribution is an embodiment-aware evaluation framework and paired pseudo-ground-truth data for scaling robot-centric policy pretraining from abundant human video, while acknowledging that synthetic composites cannot replace real-robot observations.
Original abstract
Robotic manipulation with dexterous hands is a cornerstone of Embodied AI, yet its progress is stifled by the high cost of collecting embodiment-aware teleoperation data. While abundant egocentric videos of human hands offer a scalable alternative, the profound discrepancies in appearance, articulation, and camera viewpoints between human and robotic data raise significant challenges for co-training. Though existing general image-editing models demonstrate strong capabilities, they lack necessary embodiment-specific priors to fully bridge this gap. In this work, we present HandEdit, a unified large-scale embodiment-aware image-editing dataset and benchmark specifically designed to transform human hands and arms into various dexterous robotic embodiments within egocentric frames. HandEdit comprises over 200M editing instances derived from five diverse source datasets, covering 26 distinct URDFs, including 13 hand-only and 13 hand-arm configurations. Alongside the dataset, we establish a unified benchmark protocol with two tracks: Hand-only and Hand-Arm, supporting URDF-conditioned evaluation. We conduct extensive evaluations of 11 representative image-editing baselines using a multi-dimensional metric suite, including generic similarity metrics, VLM-based judgment, and embodiment-aware metrics. HandEdit serves as a critical resource at the intersection of image editing and robotics: it advances embodiment-aware editing models while enabling scalable dexterous robotic learning from abundant human video data, paving the way for more generalizable Embodied AI.
Read the original paperMore in Embodied AI
Browse all 48 papers →GroundingPI: A Grounding Foundation Model towards Physical Intelligence with Visual Primitives
Qize Yu, Lianrui Fan, Boyu Chen, Jiaqi Liang, Xini Ding, Yue Chen, Zetian Song, Yuran Wang, Yi Zou, Kaixuan Wang, Tianxing Chen, Wenxuan Song, Bohan Zhou, Mingleyang Li, Siqiao Huang, Yuqi Ye, Caigao Jiang, Wei Wei, Ruihai Wu, Hang Zhang, Yixiao Ge, Shuchang Zhou, Shilong Liu, Xianming Liu, Ping Luo, Shiyu Huang
GroundingPI argues that fast, precise visual grounding should be the perceptual foundation for capable robots and autonomous vehicles.
MM-ABC: Towards Generalist Mobile Manipulation via Seeing, Coordinating and Imagining
Qiwei Liang, Guangyu Chen, Shaolong Zhu, Zikuan Xiao, Jinxuan Lu, Yifan Xie, Renjing Xu, Wenbo Ding, Tianxing Chen
MM-ABC is a generalist robot foundation model that helps mobile manipulators see their surroundings, coordinate arm and base motion, and imagine future actions for better performance.
Morphometric Imitation: From Morphology and Contact Aware Hand Retargeting to Sim-to-Real Visuomotor Policy
Tara Sadjadpour, Siming He, C. K. Wolfe, Haozhi Qi, Lea Wilken, S. Shankar Sastry, Claire Tomlin, Jitendra Malik
A three-stage system converts human hand demonstrations into robust, zero-shot real-robot dexterous manipulation policies across different hand morphologies.