LeVJEPA: Efficient & Scalable Video Pretraining without the Heuristics
AuthorsLukas Kuhn, Lucas Maes, Giuseppe Serra, Quentin Le Lidec, Yann LeCun, Randall Balestriero, Florian Buettner
Resources
LeVJEPA aims to make video pretraining dramatically cheaper and simpler by learning useful representations without momentum encoders, stop-gradients, or pixel reconstruction.
Key results
The sole objective hyperparameter, fixed in every experiment.
Uniform random patch-token dropping used during pretraining.
Largest reported pretraining-compute reduction across ViT-S, ViT-B, and ViT-L.
LeVJEPA ViT-B accuracy under matched compute on the 20% K710 setup.
LeVJEPA ViT-B linear-probing accuracy under matched compute.
LeVJEPA accuracy versus 16.9% for DINOv2 at matched compute.
What the paper found
LeVJEPA is a video-pretraining method that removes the usual anti-collapse heuristics used by models such as Meta’s V-JEPA 2: there is no exponential-moving-average target encoder, predictor, stop-gradient, or masked-pixel decoder. Instead, one shared video transformer encoder and a small projector optimize mean-squared invariance between global and local clip views, while SIGReg forces embeddings toward an isotropic Gaussian with a provable anti-collapse guarantee; its only loss hyperparameter is fixed at 0.02. Training uses 16-frame clips from a 20% subsample of K710 and uniformly discards 95% of patch tokens, reducing compute while improving ImageNet-1K probing from 33.9% with no dropping to 47.6%. Across ViT-S, ViT-B, and ViT-L, LeVJEPA matches or exceeds V-JEPA 2 with 5.6 to 20.8 times less pretraining compute; at matched FLOPs, it reaches 61.0% on ImageNet-1K and 44.6% on Kinetics-400, leading the compared video baselines by 7.6 points on ImageNet-1K. Block-causal attention, bidirectional within each frame but causal across time, slightly beats bidirectional attention, 51.2% versus 50.7%, while enabling frame representations that use only current and past observations. Against DINOv2 trained on frames from the same videos, LeVJEPA scores 30.4% on Something-Something-v2 versus 16.9% for DINOv2, showing that efficient video pretraining preserves much of image-level recognition while substantially improving motion understanding.
Original abstract
Video carries the temporal structure of the physical world, yet learning representations from it has remained computationally expensive: prevailing self-supervised methods either prevent representation collapse through architectural asymmetries, coupling an exponential-moving-average target encoder, a stop-gradient, and a capacity-limited predictor, or circumvent it by reconstructing masked content in pixel space. We introduce LeVJEPA, the first video encoder trained under LeJEPA's collapse-free objective, which dispenses with both. A single encoder is trained with an invariance loss over global and local views of a clip, regularized by SIGReg, which excludes collapse with a provable guarantee. The architecture reduces to an encoder and a projector, and the objective to a single hyperparameter. This formulation admits two properties. First, the cost of pretraining is governed by the number of tokens the encoder observes; uniform random token dropping renders this number small while simultaneously improving downstream accuracy. At matched epochs on identical data, LeVJEPA matches or surpasses V-JEPA 2 across ViT-S/B/L at 5.6 to 20.8x less pretraining compute, and at matched total FLOPs it exceeds the strongest video baseline by 7.6 points on ImageNet-1K while remaining competitive on motion-centric benchmarks. Second, since no asymmetry between branches is required, the encoder can be trained with block-causal attention at no measurable accuracy cost: temporal ordering becomes a property of the encoder itself. Against a compute-matched DINOv2 trained on frames of the same videos, LeVJEPA approaches the image-pretrained encoder on appearance-centric evaluation while nearly doubling its motion-centric accuracy. These results indicate that, once its computational overhead is removed, video becomes a viable and in several respects preferable substrate for general-purpose visual pretraining.
Read the original paperMore in Self-Supervised Learning
Browse all 22 papers →Self-Play Pretraining with Zero Data
Aditya Cowsik, Kfir Dolev, Michael Y. Li, G. Bruno De Luca, Nourya Cohen, Noah D. Goodman, Yoav Levine
A learner and an RL-powered program generator teach each other from scratch, producing synthetic data that enables surprisingly meaningful transfer to natural datasets.
Strategically Diverse Sampling for Self-Training
Alexander Gurung, Esmeralda S. Whitammer, Mirella Lapata
Instead of training LLMs on many similar correct answers, this work shows that exposing them to diverse problem-solving strategies—even imperfect ones—can produce stronger models.
TT-VidT: Decoupling the Temporal Axis for Efficient Motion-Centric Video Pretraining
Shih-Ying Yeh, Daniel Z. Kaplan, Xuehai Wang, Fu-En Yang, Min-Hung Chen, Shang-Hong Lai
TT-VidT pretrains video models to focus on motion while preserving appearance, achieving strong action-recognition results with substantially lower compute.