NTH

LeVJEPA: Efficient & Scalable Video Pretraining without the Heuristics

AuthorsLukas Kuhn, Lucas Maes, Giuseppe Serra, Quentin Le Lidec, Yann LeCun, Randall Balestriero, Florian Buettner

September 2, 2026 3 min read
Watch on YouTube
The one-line take

LeVJEPA aims to make video pretraining dramatically cheaper and simpler by learning useful representations without momentum encoders, stop-gradients, or pixel reconstruction.

Key results

0.02
SIGReg loss weight

The sole objective hyperparameter, fixed in every experiment.

95%
Token dropping ratio

Uniform random patch-token dropping used during pretraining.

20.8
Maximum compute reduction versus V-JEPA 2

Largest reported pretraining-compute reduction across ViT-S, ViT-B, and ViT-L.

61.0
FLOP-matched ImageNet-1K accuracy

LeVJEPA ViT-B accuracy under matched compute on the 20% K710 setup.

44.6
FLOP-matched Kinetics-400 accuracy

LeVJEPA ViT-B linear-probing accuracy under matched compute.

30.4%
Something-Something-v2 accuracy

LeVJEPA accuracy versus 16.9% for DINOv2 at matched compute.

What the paper found

LeVJEPA is a video-pretraining method that removes the usual anti-collapse heuristics used by models such as Meta’s V-JEPA 2: there is no exponential-moving-average target encoder, predictor, stop-gradient, or masked-pixel decoder. Instead, one shared video transformer encoder and a small projector optimize mean-squared invariance between global and local clip views, while SIGReg forces embeddings toward an isotropic Gaussian with a provable anti-collapse guarantee; its only loss hyperparameter is fixed at 0.02. Training uses 16-frame clips from a 20% subsample of K710 and uniformly discards 95% of patch tokens, reducing compute while improving ImageNet-1K probing from 33.9% with no dropping to 47.6%. Across ViT-S, ViT-B, and ViT-L, LeVJEPA matches or exceeds V-JEPA 2 with 5.6 to 20.8 times less pretraining compute; at matched FLOPs, it reaches 61.0% on ImageNet-1K and 44.6% on Kinetics-400, leading the compared video baselines by 7.6 points on ImageNet-1K. Block-causal attention, bidirectional within each frame but causal across time, slightly beats bidirectional attention, 51.2% versus 50.7%, while enabling frame representations that use only current and past observations. Against DINOv2 trained on frames from the same videos, LeVJEPA scores 30.4% on Something-Something-v2 versus 16.9% for DINOv2, showing that efficient video pretraining preserves much of image-level recognition while substantially improving motion understanding.

Original abstract

Video carries the temporal structure of the physical world, yet learning representations from it has remained computationally expensive: prevailing self-supervised methods either prevent representation collapse through architectural asymmetries, coupling an exponential-moving-average target encoder, a stop-gradient, and a capacity-limited predictor, or circumvent it by reconstructing masked content in pixel space. We introduce LeVJEPA, the first video encoder trained under LeJEPA's collapse-free objective, which dispenses with both. A single encoder is trained with an invariance loss over global and local views of a clip, regularized by SIGReg, which excludes collapse with a provable guarantee. The architecture reduces to an encoder and a projector, and the objective to a single hyperparameter. This formulation admits two properties. First, the cost of pretraining is governed by the number of tokens the encoder observes; uniform random token dropping renders this number small while simultaneously improving downstream accuracy. At matched epochs on identical data, LeVJEPA matches or surpasses V-JEPA 2 across ViT-S/B/L at 5.6 to 20.8x less pretraining compute, and at matched total FLOPs it exceeds the strongest video baseline by 7.6 points on ImageNet-1K while remaining competitive on motion-centric benchmarks. Second, since no asymmetry between branches is required, the encoder can be trained with block-causal attention at no measurable accuracy cost: temporal ordering becomes a property of the encoder itself. Against a compute-matched DINOv2 trained on frames of the same videos, LeVJEPA approaches the image-pretrained encoder on appearance-centric evaluation while nearly doubling its motion-centric accuracy. These results indicate that, once its computational overhead is removed, video becomes a viable and in several respects preferable substrate for general-purpose visual pretraining.

Read the original paper

More in Self-Supervised Learning

Browse all 22 papers →
01Self Supervised

Self-Play Pretraining with Zero Data

Aditya Cowsik, Kfir Dolev, Michael Y. Li, G. Bruno De Luca, Nourya Cohen, Noah D. Goodman, Yoav Levine

A learner and an RL-powered program generator teach each other from scratch, producing synthetic data that enables surprisingly meaningful transfer to natural datasets.

Read analysis
02Self Supervised

Strategically Diverse Sampling for Self-Training

Alexander Gurung, Esmeralda S. Whitammer, Mirella Lapata

Instead of training LLMs on many similar correct answers, this work shows that exposing them to diverse problem-solving strategies—even imperfect ones—can produce stronger models.

Read analysis