The Surprising Effectiveness of Video Diffusion Models for Hand Motion Reconstruction
AuthorsYuxi Wang, Chengkai Jin, Yufei Liu, Wenqi Ouyang, Tianyi Wei, Zhiwei Zeng, Siyuan Huang, Zhiqi Shen, Xingang Pan
Resources
This work shows that video diffusion models can do more than generate clips—they can also reconstruct detailed two-hand motion from egocentric video surprisingly well.
Key results
Wan2.1-VACE video diffusion transformer used as the pretrained backbone
Dual-branch hand decoder parameter count
Mid-block feature slice used for decoding
Best flow-matching time for feature extraction
Best frame accuracy on ARCTIC
Penalty MPJPE on ARCTIC in mm
What the paper found
ViDiHand is a new hand-motion reconstruction system that repurposes the 1.3B-parameter Wan2.1-VACE video diffusion model, an Alibaba Group backbone, for 4D two-hand pose recovery from egocentric video. The key idea is to finetune only the VACE branch with a hand-overlay rendering objective, first on EgoDex joint overlays and then on ARCTIC and HOT3D MANO mesh overlays, so the diffusion model learns hand-aware scene features without losing its general video prior. A 37M-parameter dual-branch decoder then reads a single mid-layer activation, specifically DiT layer 15 at denoising step 0.7, and combines a hand-token branch, a 21-joint heatmap branch, mutual cross-attention, and a mixed-projection camera-translation solve to recover MANO pose and metric-scale 3D translation with no detector, infiller, or test-time optimization. On ARCTIC, ViDiHand reaches 0.997 frame accuracy, 21.668 mm MPJPE-p, 12.407 px EPE-p, and 3.183 mm/frame2 jitter; on HOT3D it reaches 0.948 FAcc and 3.741 jitter; on held-out HOI4D it reaches 0.984 FAcc and 4.010 jitter, ranking first on eight of nine metrics. Ablations show that the full decoder improves ARCTIC FAcc from 0.9767–0.9829 in reduced variants to 0.9979, and removing acceleration smoothness raises jitter from 3.42 to 3.88, confirming that the diffusion prior supplies both occlusion robustness and temporal stability.
Original abstract
4D hand motion reconstruction from egocentric video is bottlenecked by clear limitations of existing methods: image-based pipelines depend on a detector that fails under heavy occlusion, while video-based methods rely on temporal modules learned only from scarce hand-pose annotations, a narrow signal insufficient to model motion dynamics, occlusion reasoning, and hand-object interaction. These capabilities, however, are exactly what video generative models must implicitly acquire when trained to synthesize coherent video at internet scale. Motivated by this, we present ViDiHand, which leverages the representations of a pretrained video diffusion model to reconstruct 4D two-hand pose. We adapt it via a hand-overlay rendering objective that specializes its features for hands while preserving its world priors. A decoder then recovers metric-scale pose from the adapted features. The whole pipeline operates directly on full frames--no detector, no infiller, and no test-time optimization. On ARCTIC, HOT3D, and HOI4D, ViDiHand substantially outperforms prior methods, establishing video diffusion models as a powerful new foundation for hand motion reconstruction and a promising route to scalable in-the-wild data collection for embodied AI. Project page: https://vidihand.github.io.
Read the original paperMore in Computer Vision
Browse all 58 papers →All-in-One Multilingual Scene Text Recognition with Script-aware Mixture-of-Experts
Xingsong Ye, Yongkun Du, Jiaxin Zhang, Zhixian Li, Chong Sun, Chen Li, Jing Lyu, Lianwen Jin, Zhineng Chen
A lightweight script-aware mixture-of-experts model brings more accurate, scalable multilingual scene text recognition to many languages and scripts.
DyRAD: Radar Novel View Synthesis for Dynamic Driving Scenes
Merav Keidar, Tomer Borreda, Rajalakshmi Nandakumar, Or Litany
DyRAD builds moving radar views of driving scenes by combining tracked object motion with the radar’s physics, enabling more realistic and transferable autonomy testing.
OmniTaskonomy: When Does Visual Generation Improve Visual Understanding?
Jiaxin Ge, Yiming Qin, Ji Xie, Haozhe Jiang, Xiaochuang Han, Junyi Zhang, Andrew Dai, Yinfei Yang, Jitendra Malik, Ranjay Krishna, Sewon Min, Haiwen Feng, Le Xue, Baifeng Shi, Trevor Darrell, XuDong Wang
This work maps when training models to generate images can make them better at understanding images, revealing both intuitive and surprising task-to-task benefits.