NTH
Research collection

Self-Supervised Learning research

Explore methods that learn representations from data without task-specific labels. Follow research on training objectives, transfer, and data efficiency.

22 papers · Latest edition September 30, 2026

Where to start

Three of the latest briefs in this collection. Read the evidence and the original papers alongside them.

All Self-Supervised Learning papers

Newest editions first.

01Self Supervised

Self-Play Pretraining with Zero Data

Aditya Cowsik, Kfir Dolev, Michael Y. Li, G. Bruno De Luca, Nourya Cohen, Noah D. Goodman, Yoav Levine

A learner and an RL-powered program generator teach each other from scratch, producing synthetic data that enables surprisingly meaningful transfer to natural datasets.

Read analysis
02Self Supervised

Strategically Diverse Sampling for Self-Training

Alexander Gurung, Esmeralda S. Whitammer, Mirella Lapata

Instead of training LLMs on many similar correct answers, this work shows that exposing them to diverse problem-solving strategies—even imperfect ones—can produce stronger models.

Read analysis
05Self Supervised

LeVJEPA: Efficient & Scalable Video Pretraining without the Heuristics

Lukas Kuhn, Lucas Maes, Giuseppe Serra, Quentin Le Lidec, Yann LeCun, Randall Balestriero, Florian Buettner

LeVJEPA aims to make video pretraining dramatically cheaper and simpler by learning useful representations without momentum encoders, stop-gradients, or pixel reconstruction.

Read analysis
06Self Supervised

TTPO: Test-Time Policy Optimization

Aozhe Wang, Zhengxi Lu, Jianze Wang, Shangke Lv, Ying Liu, Weiming Lu, Jun Xiao, Yueting Zhuang, Hua Yang, Qianglong Chen, Yongliang Shen

TTPO lets language models improve their reasoning at test time by learning from their own votes while selectively correcting likely mistakes.

Read analysis
08Self Supervised

On-Policy Self-Distillation without Any Supervision

Yijiang Li, Bingyang Wang, Yijun Liang, Yunjie Tian, Di Fu, Nuno Vasconcelos

U-OPSD lets an LLM improve its mathematical reasoning by learning from the agreement and mistakes in its own sampled solutions, without external labels or teacher models.

Read analysis
10Self Supervised

A Theory of Contrastive Learning with Natural Images

Antonio Torralba, Yair Weiss

This paper explains why contrastive learning on images works by proving that the best features are often sinusoidal filters with partial whitening, and showing real models learn the same pattern.

Read analysis
12Self Supervised

Autodata: An agentic data scientist to create high quality synthetic data

Ilia Kulikov, Chenxi Whitehouse, Tianhao Wu, Yixin Nie, Swarnadeep Saha, Eryk Helenowski, Weizhe Yuan, Olga Golovneva, Jack Lanchantin, Yoram Bachrach, Jakob Foerster, Xian Li, Han Fang, Sainbayar Sukhbaatar, Jason Weston

Autodata turns an AI agent into a data scientist that generates better training data, and even learns how to improve its own data-making process.

Read analysis
13Self Supervised

Vision Pretraining for Dense Spatial Perception

Zelin Fu, Bin Tan, Changjiang Sun, Shaohui Liu, Kecheng Zheng, Yinghao Xu, Xing Zhu, Yujun Shen, Nan Xue

This paper teaches vision models to pay attention to boundaries so they can learn richer spatial structure and improve tasks like depth estimation.

Read analysis
14Self Supervised

Toward Parking Spot Occupancy Recognition: A Self-Supervised Approach

Luan Marko Kujavski, Rayson Laroca, Paulo Lisboa de Almeida

This paper shows that self-supervised learning can remove the need for target-lot labels while still achieving very high parking-spot occupancy accuracy, making deployment cheaper and more scalable.

Read analysis