A Theory of Contrastive Learning with Natural Images
AuthorsAntonio Torralba, Yair Weiss
Resources
This paper explains why contrastive learning on images works by proving that the best features are often sinusoidal filters with partial whitening, and showing real models learn the same pattern.
Key results
Best transfer from synthetic training distributions to CIFAR10
Linear partial whitening baseline on CIFAR10
What the paper found
Antonio Torralba and Yair Weiss show that contrastive learning on natural images can be understood analytically through frequency-domain statistics rather than heuristic invariances. By replacing InfoNCE with a Gaussian-equivalent objective, they prove that the optimum is a “partial whitening” representation: it selects a subset of Fourier modes and rescales them inversely to expected power, so the embedding covariance becomes white. For simple augmentations such as circular crop, brightness/contrast jitter, and ideal blur, the globally optimal representation is just squared DFT magnitudes at K non-DC frequencies, normalized to unit norm; this can be computed by a one-hidden-layer CNN with sinusoidal filters, pointwise nonlinearity, global average pooling, and a linear projection. For harder augmentations including crop-plus-noise, arbitrary blur, and linear jitter, the optimal last-layer weights are generalized eigenvectors of the alignment matrix B and covariance Σ, with the selected frequencies found by a waterfilling algorithm. The theory predicts, and experiments confirm, that SGD-trained CNNs on CIFAR10, CIFAR100, ImageNet, fractal noise, dead leaves, and CelebA learn sinusoidal first-layer filters and frequency sensitivities shaped like rings or diamonds in Fourier space. In their CIFAR10 KNN study, training on non-CIFAR synthetic noise yields the best transfer when its expected power spectrum matches CIFAR10, reaching 45.1%, while simple partial whitening in PCA space gives 47.1% and close to the learned representation’s performance.
Original abstract
Why does contrastive learning with simple images and augmentations yield useful representations for downstream tasks? We address this question by analytically computing the optimal representation in terms of a contrastive loss for a range of basic augmentations and any image dataset with stationary statistics. We show that for certain augmentations the optimum can be attained by a CNN whose first layer filters are sinusoids, followed by a pointwise nonlinearity, global average pooling, and a final linear layer that performs partial whitening. We also show that the optimal weights in such CNNs for more complicated augmentations are still sinusoids. The frequencies of the sinusoids and their weights can be computed using a simple waterfilling algorithm given the dataset's expected power spectrum. Experiments with different image datasets and augmentations show that such CNNs trained with SGD empirically learn sinusoids in their first layer and to perform partial whitening
Read the original paperMore in Self-Supervised Learning
Browse all 22 papers →Self-Play Pretraining with Zero Data
Aditya Cowsik, Kfir Dolev, Michael Y. Li, G. Bruno De Luca, Nourya Cohen, Noah D. Goodman, Yoav Levine
A learner and an RL-powered program generator teach each other from scratch, producing synthetic data that enables surprisingly meaningful transfer to natural datasets.
Strategically Diverse Sampling for Self-Training
Alexander Gurung, Esmeralda S. Whitammer, Mirella Lapata
Instead of training LLMs on many similar correct answers, this work shows that exposing them to diverse problem-solving strategies—even imperfect ones—can produce stronger models.
TT-VidT: Decoupling the Temporal Axis for Efficient Motion-Centric Video Pretraining
Shih-Ying Yeh, Daniel Z. Kaplan, Xuehai Wang, Fu-En Yang, Min-Hung Chen, Shang-Hong Lai
TT-VidT pretrains video models to focus on motion while preserving appearance, achieving strong action-recognition results with substantially lower compute.