Explore methods that learn representations from data without task-specific labels. Follow research on training objectives, transfer, and data efficiency.
22 papers · Latest edition September 30, 2026
Where to start
Three of the latest briefs in this collection. Read the evidence and the original papers alongside them.
A learner and an RL-powered program generator teach each other from scratch, producing synthetic data that enables surprisingly meaningful transfer to natural datasets.
Instead of training LLMs on many similar correct answers, this work shows that exposing them to diverse problem-solving strategies—even imperfect ones—can produce stronger models.
TT-VidT pretrains video models to focus on motion while preserving appearance, achieving strong action-recognition results with substantially lower compute.
Aditya Cowsik, Kfir Dolev, Michael Y. Li, G. Bruno De Luca, Nourya Cohen, Noah D. Goodman, Yoav Levine
A learner and an RL-powered program generator teach each other from scratch, producing synthetic data that enables surprisingly meaningful transfer to natural datasets.
Alexander Gurung, Esmeralda S. Whitammer, Mirella Lapata
Instead of training LLMs on many similar correct answers, this work shows that exposing them to diverse problem-solving strategies—even imperfect ones—can produce stronger models.
Shih-Ying Yeh, Daniel Z. Kaplan, Xuehai Wang, Fu-En Yang, Min-Hung Chen, Shang-Hong Lai
TT-VidT pretrains video models to focus on motion while preserving appearance, achieving strong action-recognition results with substantially lower compute.
Lukas Kuhn, Lucas Maes, Giuseppe Serra, Quentin Le Lidec, Yann LeCun, Randall Balestriero, Florian Buettner
LeVJEPA aims to make video pretraining dramatically cheaper and simpler by learning useful representations without momentum encoders, stop-gradients, or pixel reconstruction.
Yongkang Yang, Zhezheng Hao, Hong Zhang, Yi Liu, Xiankun Lin, Wence Ji, Fanjunduo Wei, Jiarui Yu, Qiang Lin, Xiaoyun Liang, Hande Dong
USD teaches language models to choose both what supervision to absorb and how difficult that supervision should be, adapting self-distillation to the model’s changing learning capacity.
U-OPSD lets an LLM improve its mathematical reasoning by learning from the agreement and mistakes in its own sampled solutions, without external labels or teacher models.
CVPD helps multimodal models notice visual details they can perceive but previously failed to use, without relying on external teachers or annotations.
This paper explains why contrastive learning on images works by proving that the best features are often sinusoidal filters with partial whitening, and showing real models learn the same pattern.
SiamJEPA shows that giving JEPA a Siamese student encoder can act like a useful regularizer, improving early learning and representation quality in self-supervised vision models.
Ilia Kulikov, Chenxi Whitehouse, Tianhao Wu, Yixin Nie, Swarnadeep Saha, Eryk Helenowski, Weizhe Yuan, Olga Golovneva, Jack Lanchantin, Yoram Bachrach, Jakob Foerster, Xian Li, Han Fang, Sainbayar Sukhbaatar, Jason Weston
Autodata turns an AI agent into a data scientist that generates better training data, and even learns how to improve its own data-making process.
Luan Marko Kujavski, Rayson Laroca, Paulo Lisboa de Almeida
This paper shows that self-supervised learning can remove the need for target-lot labels while still achieving very high parking-spot occupancy accuracy, making deployment cheaper and more scalable.
Yohann Perron, Guillaume Astruc, Nicolas Gonthier, Clement Mallet, Loic Landrieu
UniverSat makes one transformer backbone work across many Earth-observation sensors, resolutions, and modalities by learning a shared patch embedding through self-supervision.
This paper improves self-supervised learning on medical tabular data by adaptively refining feature discretization during training, leading to better representations without hand-tuned binning.
U-TTT adapts a PET denoising model on the fly during inference, using self-supervised spatial and frequency updates to stay robust across scanners and dose shifts.
Ninad Daithankar, Alexi Gladstone, Yann LeCun, Heng Ji
This paper introduces a new way to learn visual representations from video by predicting how representations change over time, potentially reducing reliance on hand-crafted training tricks.
Chunzheng Zhu, Yijun Wang, Jianxin Lin, Feng Wang, Hongwei Wang, Lei Zhao, Shengli Li, Kenli Li
This paper teaches ultrasound models to learn from anatomy, not just pixels, by using self-supervision tied to clinically meaningful structures for more robust medical imaging representations.
This paper introduces a faster, augmentation-free way to learn strong time-series representations by contrasting meaningful chunks of each sequence instead of individual timesteps.
This paper claims many common robustness tricks are really different ways of estimating the same nuisance-covariance object, and builds a geometric theory showing how to match regularization to that nuisance structure.