NTH

VideoMDM: Towards 3D Human Motion Generation From 2D Supervision

AuthorsAmir Mann, Gal Michael Harari, Merav Keidar, Or Litany

July 12, 2026 2 min read
Watch on YouTube
The one-line take

VideoMDM teaches a 3D human motion generator using only 2D keypoints from videos, getting surprisingly close to fully 3D-supervised models without needing 3D ground truth.

Key results

0.88
HumanML3D FID

VideoMDM on the 2D-only HumanML3D setup

0.54
MDM FID

Fully 3D-supervised MDM on HumanML3D

14,616
HumanML3D sequences

Size of the HumanML3D motion dataset used in evaluation

111.24
Fit3D MPJPE

VideoMDM on Fit3D lifting evaluation, in mm

228.47
WHAM MPJPE

Baseline on Fit3D, in mm

64.0%
NBA human preference

Pairwise human preference for VideoMDM over MAS

What the paper found

VideoMDM, developed by Amir Mann, Gal Michael Harari, Merav Keidar, and Or Litany from Technion and NVIDIA, reframes 3D human motion generation as a cross-modality diffusion problem trained only on 2D pose supervision from monocular video. A pretrained 2D-to-3D lifter supplies noisy 3D teacher trajectories, but the model learns natively in 3D and is supervised through depth-aware 2D reprojection, a loss the authors show is equivalent in expectation to direct 3D MSE under mild camera assumptions. The method also adapts 3D motion regularizers to 2D, including a depth-weighted velocity loss and a ray-projection-based representation alignment term for rotations, joint velocities, and foot contacts. On a 2D-only HumanML3D setup with 14,616 motion sequences, VideoMDM reaches FID 0.88, nearly closing the gap to fully 3D-supervised MDM at 0.54 and beating the strongest 2D baseline. On Fit3D, it cuts MPJPE from 228.47 mm for WHAM to 111.24 mm and reduces acceleration error from 17.66 to 3.16 m/s². On NBA, human raters preferred VideoMDM over MAS in 64.0% of pairwise comparisons, showing that 2D supervision alone can produce a coherent 3D motion prior that generalizes beyond the lifter that bootstrapped it.

Original abstract

We introduce VideoMDM, a diffusion-based framework that trains 3D human motion priors directly from accurate 2D poses extracted from monocular videos, without any 3D ground truth. A pretrained 2D-to-3D lifter provides approximate 3D pose sequences that serve as a noisy teacher: these are diffused, denoised by the model in 3D, and supervised in 2D by reprojecting the prediction and comparing against accurate keypoints. We show that, under mild assumptions, a depth-weighted 2D reprojection loss is equivalent in expectation to direct 3D supervision, and we adapt standard 3D motion regularizers - velocity consistency and over-parameterized representation alignment - to this 2D setting. Unlike methods that lift 2D to 3D only at inference, VideoMDM learns a coherent 3D motion manifold during training. On HumanML3D it nearly closes the gap to fully 3D-supervised MDM (FID 0.88 vs 0.54); On real video datasets Fit3D and NBA the method learns to generate motions consistently preferred by humans, with strong quantitative results.

Read the original paper

More in Computer Vision

Browse all 58 papers →
02Cv

DyRAD: Radar Novel View Synthesis for Dynamic Driving Scenes

Merav Keidar, Tomer Borreda, Rajalakshmi Nandakumar, Or Litany

DyRAD builds moving radar views of driving scenes by combining tracked object motion with the radar’s physics, enabling more realistic and transferable autonomy testing.

Read analysis
03Cv

OmniTaskonomy: When Does Visual Generation Improve Visual Understanding?

Jiaxin Ge, Yiming Qin, Ji Xie, Haozhe Jiang, Xiaochuang Han, Junyi Zhang, Andrew Dai, Yinfei Yang, Jitendra Malik, Ranjay Krishna, Sewon Min, Haiwen Feng, Le Xue, Baifeng Shi, Trevor Darrell, XuDong Wang

This work maps when training models to generate images can make them better at understanding images, revealing both intuitive and surprising task-to-task benefits.

Read analysis