VideoMDM: Towards 3D Human Motion Generation From 2D Supervision
AuthorsAmir Mann, Gal Michael Harari, Merav Keidar, Or Litany
Resources
VideoMDM teaches a 3D human motion generator using only 2D keypoints from videos, getting surprisingly close to fully 3D-supervised models without needing 3D ground truth.
Key results
VideoMDM on the 2D-only HumanML3D setup
Fully 3D-supervised MDM on HumanML3D
Size of the HumanML3D motion dataset used in evaluation
VideoMDM on Fit3D lifting evaluation, in mm
Baseline on Fit3D, in mm
Pairwise human preference for VideoMDM over MAS
What the paper found
VideoMDM, developed by Amir Mann, Gal Michael Harari, Merav Keidar, and Or Litany from Technion and NVIDIA, reframes 3D human motion generation as a cross-modality diffusion problem trained only on 2D pose supervision from monocular video. A pretrained 2D-to-3D lifter supplies noisy 3D teacher trajectories, but the model learns natively in 3D and is supervised through depth-aware 2D reprojection, a loss the authors show is equivalent in expectation to direct 3D MSE under mild camera assumptions. The method also adapts 3D motion regularizers to 2D, including a depth-weighted velocity loss and a ray-projection-based representation alignment term for rotations, joint velocities, and foot contacts. On a 2D-only HumanML3D setup with 14,616 motion sequences, VideoMDM reaches FID 0.88, nearly closing the gap to fully 3D-supervised MDM at 0.54 and beating the strongest 2D baseline. On Fit3D, it cuts MPJPE from 228.47 mm for WHAM to 111.24 mm and reduces acceleration error from 17.66 to 3.16 m/s². On NBA, human raters preferred VideoMDM over MAS in 64.0% of pairwise comparisons, showing that 2D supervision alone can produce a coherent 3D motion prior that generalizes beyond the lifter that bootstrapped it.
Original abstract
We introduce VideoMDM, a diffusion-based framework that trains 3D human motion priors directly from accurate 2D poses extracted from monocular videos, without any 3D ground truth. A pretrained 2D-to-3D lifter provides approximate 3D pose sequences that serve as a noisy teacher: these are diffused, denoised by the model in 3D, and supervised in 2D by reprojecting the prediction and comparing against accurate keypoints. We show that, under mild assumptions, a depth-weighted 2D reprojection loss is equivalent in expectation to direct 3D supervision, and we adapt standard 3D motion regularizers - velocity consistency and over-parameterized representation alignment - to this 2D setting. Unlike methods that lift 2D to 3D only at inference, VideoMDM learns a coherent 3D motion manifold during training. On HumanML3D it nearly closes the gap to fully 3D-supervised MDM (FID 0.88 vs 0.54); On real video datasets Fit3D and NBA the method learns to generate motions consistently preferred by humans, with strong quantitative results.
Read the original paperMore in Computer Vision
Browse all 58 papers →All-in-One Multilingual Scene Text Recognition with Script-aware Mixture-of-Experts
Xingsong Ye, Yongkun Du, Jiaxin Zhang, Zhixian Li, Chong Sun, Chen Li, Jing Lyu, Lianwen Jin, Zhineng Chen
A lightweight script-aware mixture-of-experts model brings more accurate, scalable multilingual scene text recognition to many languages and scripts.
DyRAD: Radar Novel View Synthesis for Dynamic Driving Scenes
Merav Keidar, Tomer Borreda, Rajalakshmi Nandakumar, Or Litany
DyRAD builds moving radar views of driving scenes by combining tracked object motion with the radar’s physics, enabling more realistic and transferable autonomy testing.
OmniTaskonomy: When Does Visual Generation Improve Visual Understanding?
Jiaxin Ge, Yiming Qin, Ji Xie, Haozhe Jiang, Xiaochuang Han, Junyi Zhang, Andrew Dai, Yinfei Yang, Jitendra Malik, Ranjay Krishna, Sewon Min, Haiwen Feng, Le Xue, Baifeng Shi, Trevor Darrell, XuDong Wang
This work maps when training models to generate images can make them better at understanding images, revealing both intuitive and surprising task-to-task benefits.