NTH

A Theory of Contrastive Learning with Natural Images

AuthorsAntonio Torralba, Yair Weiss

July 10, 2026 2 min read
Watch on YouTube
The one-line take

This paper explains why contrastive learning on images works by proving that the best features are often sinusoidal filters with partial whitening, and showing real models learn the same pattern.

Key results

45.1%
Synthetic-to-CIFAR10 KNN

Best transfer from synthetic training distributions to CIFAR10

47.1%
PCA partial whitening KNN

Linear partial whitening baseline on CIFAR10

What the paper found

Antonio Torralba and Yair Weiss show that contrastive learning on natural images can be understood analytically through frequency-domain statistics rather than heuristic invariances. By replacing InfoNCE with a Gaussian-equivalent objective, they prove that the optimum is a “partial whitening” representation: it selects a subset of Fourier modes and rescales them inversely to expected power, so the embedding covariance becomes white. For simple augmentations such as circular crop, brightness/contrast jitter, and ideal blur, the globally optimal representation is just squared DFT magnitudes at K non-DC frequencies, normalized to unit norm; this can be computed by a one-hidden-layer CNN with sinusoidal filters, pointwise nonlinearity, global average pooling, and a linear projection. For harder augmentations including crop-plus-noise, arbitrary blur, and linear jitter, the optimal last-layer weights are generalized eigenvectors of the alignment matrix B and covariance Σ, with the selected frequencies found by a waterfilling algorithm. The theory predicts, and experiments confirm, that SGD-trained CNNs on CIFAR10, CIFAR100, ImageNet, fractal noise, dead leaves, and CelebA learn sinusoidal first-layer filters and frequency sensitivities shaped like rings or diamonds in Fourier space. In their CIFAR10 KNN study, training on non-CIFAR synthetic noise yields the best transfer when its expected power spectrum matches CIFAR10, reaching 45.1%, while simple partial whitening in PCA space gives 47.1% and close to the learned representation’s performance.

Original abstract

Why does contrastive learning with simple images and augmentations yield useful representations for downstream tasks? We address this question by analytically computing the optimal representation in terms of a contrastive loss for a range of basic augmentations and any image dataset with stationary statistics. We show that for certain augmentations the optimum can be attained by a CNN whose first layer filters are sinusoids, followed by a pointwise nonlinearity, global average pooling, and a final linear layer that performs partial whitening. We also show that the optimal weights in such CNNs for more complicated augmentations are still sinusoids. The frequencies of the sinusoids and their weights can be computed using a simple waterfilling algorithm given the dataset's expected power spectrum. Experiments with different image datasets and augmentations show that such CNNs trained with SGD empirically learn sinusoids in their first layer and to perform partial whitening

Read the original paper

More in Self-Supervised Learning

Browse all 22 papers →
01Self Supervised

Self-Play Pretraining with Zero Data

Aditya Cowsik, Kfir Dolev, Michael Y. Li, G. Bruno De Luca, Nourya Cohen, Noah D. Goodman, Yoav Levine

A learner and an RL-powered program generator teach each other from scratch, producing synthetic data that enables surprisingly meaningful transfer to natural datasets.

Read analysis
02Self Supervised

Strategically Diverse Sampling for Self-Training

Alexander Gurung, Esmeralda S. Whitammer, Mirella Lapata

Instead of training LLMs on many similar correct answers, this work shows that exposing them to diverse problem-solving strategies—even imperfect ones—can produce stronger models.

Read analysis