NTH

Human-JEPA: A Human-Centric Vision Model that Perceives and Anticipates

AuthorsHui Wei, Licai Sun, Guoying Zhao

September 2, 2026 2 min read
Watch on YouTube
The one-line take

Human-JEPA aims to help vision systems understand people now and predict what they will do next using one efficient self-supervised video model.

Key results

536,699
Kinetics-700 training clips

Human-JEPA video pretraining corpus

958k
LUPerson-T image crops

Image branch used to preserve appearance and identity features

0.620
Pose AP

Frozen COCO keypoint probe score

0.4635
Market-1501 ReID mAP

Person re-identification score

0.873
Future latent agreement

Cosine agreement with actual future latents on held-out clips

What the paper found

Human-JEPA is a human-centric video model designed to recognize the present and anticipate near-future motion, extending V-JEPA 2.1 with anchored forecasting rather than changing its architecture. Its dense context targets are tied to a frozen copy of the original initialization, preventing silent degradation during continued pretraining, while an image branch trained on LUPerson-T crops preserves appearance and identity features. The model also replaces spatial block masking with a strict past-to-future split, forcing latent prediction to represent scene dynamics instead of copying visible content. Trained on 536,699 Kinetics-700 clips plus 958k LUPerson-T crops, the 0.3B-parameter Human-JEPA reaches 0.620 AP for pose and 0.4635 mAP on Market-1501 person re-identification, outperforming the larger Sapiens2 specialist on both measures while conceding high-resolution parsing. Its predictor achieves 0.873 cosine agreement with the actual future on held-out clips, and unlike the released V-JEPA 2.1 head, it does not reduce early-action accuracy. Ablations show why the design matters: block masking causes a 17-point re-identification collapse and roughly a five-point action penalty, whereas forecasting preserves action performance near the frozen base. The paper also tests nine person-level prediction objectives and finds no reliable partner modeling, concluding that scene-level future prediction is more effective than masking or forecasting individual people.

Original abstract

Machines that understand humans should perceive the present and anticipate the future. Existing human-centric vision model are pretrained on human images, set the state of the art in static dense perception, so motion and anticipation are out of reach. Here we present Human-JEPA, a human-centric vision model trained on video by anchored forecasting: dense targets are pinned to a frozen copy of the initialization, preventing a silent collapse of dense perception, and block masks are replaced by a pure past-to-future split, avoiding a five-point action tax and a seventeen-point re-identification collapse. Under frozen probes, Human-JEPA leads the pixel-anchored specialists on pose and person re-identification at 2.7 times fewer parameters, conceding high-resolution dense parsing, and its released predictor head is the first that does not degrade anticipation. A single safely adapted model thus serves both halves of understanding humans.

Read the original paper

More in Self-Supervised Learning

Browse all 22 papers →
01Self Supervised

Self-Play Pretraining with Zero Data

Aditya Cowsik, Kfir Dolev, Michael Y. Li, G. Bruno De Luca, Nourya Cohen, Noah D. Goodman, Yoav Levine

A learner and an RL-powered program generator teach each other from scratch, producing synthetic data that enables surprisingly meaningful transfer to natural datasets.

Read analysis
02Self Supervised

Strategically Diverse Sampling for Self-Training

Alexander Gurung, Esmeralda S. Whitammer, Mirella Lapata

Instead of training LLMs on many similar correct answers, this work shows that exposing them to diverse problem-solving strategies—even imperfect ones—can produce stronger models.

Read analysis