NTH

You Don't Need Strong Assumptions: Visual Representation Learning via Temporal Differences

AuthorsNinad Daithankar, Alexi Gladstone, Yann LeCun, Heng Ji

June 17, 2026 2 min read
Watch on YouTube
The one-line take

This paper introduces a new way to learn visual representations from video by predicting how representations change over time, potentially reducing reliance on hand-crafted training tricks.

Key results

0.1%
Data-scale sweep

Smallest ImageNet-1k subset used to test how optimal inductive bias changes with scale

100%
Data-scale sweep

Full ImageNet-1k scale used in the masking-ratio experiment

20
Pretraining epochs

TDV, DINO, and iBOT pretraining length on Something-Something V2

17.05
SSv2 KNN Top-5

TDV ViT-B online ImageNet-1k KNN Top-5 after pretraining on Something-Something V2

10.97
MPI-Sintel EPE clean

TDV ViT-B optical-flow endpoint error on MPI-Sintel clean

37.33
SceneFlow bad@1px

TDV ViT-B stereo depth bad-pixel rate at 1px

What the paper found

In “You Don’t Need Strong Assumptions: Visual Representation Learning via Temporal Differences,” Ninad Daithankar, Alexi Gladstone, Yann LeCun, and Heng Ji argue that modern vision pretraining still depends too heavily on augmentations, masking, and cropping, and they introduce Temporal Difference in Vision, or TDV, a self-supervised video method that instead uses a causal next-frame constraint. TDV trains a frame encoder and a motion encoder so the current frame embedding plus an encoded RGB difference predicts the next frame embedding, with an EMA teacher and a DINO-style collapse-prevention loss. The paper’s key empirical claim is that weaker assumptions become better as scale grows: on ImageNet-1k subsets from 0.1% to 100%, the preferred masking ratio shifts from 50% toward 30% and then 10%, indicating that stronger inductive bias becomes a bottleneck at larger data scales. Pretrained on Something-Something V2 for 20 epochs, TDV reaches ImageNet KNN Top-5 of 17.05 on ViT-B, matches or slightly trails DINO and iBOT on ADE20K and Cityscapes segmentation, and is stronger on motion-centric tasks, cutting MPI-Sintel optical-flow error to 10.97 clean and 11.85 final on ViT-B while also improving SceneFlow bad-pixel rates to 54.62 at 0.5px and 37.33 at 1px. Ablations show the method depends critically on both the motion encoder and the MSE temporal-prediction loss, since removing either causes collapse.

Original abstract

Progress in AI has largely been driven by methods that assume less. As compute and data increase, approaches with weaker inductive biases generally outperform those with stronger assumptions. This is particularly characteristic of the field of Visual Representation Learning, where approaches have gone from being dominated by Supervised Learning, to Weakly Supervised Learning, to the now widespread success of Self-Supervised Learning without human labels. Yet, even modern Self-Supervised Learning approaches still depend on strong inductive biases such as augmentations, masking, or cropping. If this trend holds, even these remaining biases should become bottlenecks at scale -- and our experiments confirm this: the optimal strength of inductive biases decreases as data grows. This motivates the search for approaches that rely on fewer assumptions. To this end, we introduce Temporal Difference in Vision (TDV), a new paradigm for self-supervised learning from video that avoids existing inductive biases, relying instead on a causal assumption that the past causes the future. TDV functions by jointly training an image encoder and a motion encoder so that the current frame's representation plus the encoded motion equals the next frame's representation. Despite not leveraging any strong inductive biases, TDV matches state-of-the-art recipes on dense spatial tasks, laying the foundation for representation learning without strong assumptions.

Read the original paper

More in Self-Supervised Learning

Browse all 22 papers →
01Self Supervised

Self-Play Pretraining with Zero Data

Aditya Cowsik, Kfir Dolev, Michael Y. Li, G. Bruno De Luca, Nourya Cohen, Noah D. Goodman, Yoav Levine

A learner and an RL-powered program generator teach each other from scratch, producing synthetic data that enables surprisingly meaningful transfer to natural datasets.

Read analysis
02Self Supervised

Strategically Diverse Sampling for Self-Training

Alexander Gurung, Esmeralda S. Whitammer, Mirella Lapata

Instead of training LLMs on many similar correct answers, this work shows that exposing them to diverse problem-solving strategies—even imperfect ones—can produce stronger models.

Read analysis