NTH

TT-VidT: Decoupling the Temporal Axis for Efficient Motion-Centric Video Pretraining

AuthorsShih-Ying Yeh, Daniel Z. Kaplan, Xuehai Wang, Fu-En Yang, Min-Hung Chen, Shang-Hong Lai

AffiliationsNational Tsing Hua University Comfy Org Research realiz.ai · Karolinska Institutet Stockholm University NVIDIA†Corresponding author: kohaku@kblueleaf.netProject page: https://kohakublueleaf.github.io/TTVidT/

September 30, 2026 2 min read
Watch on YouTube
The one-line take

TT-VidT pretrains video models to focus on motion while preserving appearance, achieving strong action-recognition results with substantially lower compute.

Key results

24
Architecture-objective configurations

Matched sweep covering four encoders and six objectives

1.7M
Pretraining clip mixture

Approximate OpenVid and Moments-in-Time v2 clips

8
Pretraining duration

Epochs under the shared training recipe

73.25
Jester accuracy

TT-VidT frozen attentive-probe score

25.92
Something-Something V2 accuracy

TT-VidT frozen attentive-probe score

456.1
TT3D encoder FLOPs

Encoder-only computation in GFLOPs

What the paper found

TT-VidT targets motion-centric video representations by separating appearance from temporal change during self-supervised pretraining. Its encoder combines a DINOv3-initialized ViT-B/16 spatial path with a compact Temporal Transfer Layer, while Diff Compression reconstructs each target frame from a wide first-frame appearance anchor and only 8 frame-specific motion tokens. The study evaluates 24 architecture-objective combinations using approximately 1.7M clips from OpenVid and Moments-in-Time v2 over 8 epochs, isolating the effect of architecture and objective under a matched recipe. The strongest pairing is TT3D with Diff Compression: unlike either component alone, it reaches 73.25 on Jester, 25.92 on Something-Something V2, 37.47 on ARID, and 18.63 after Diving48 fine-tuning, leading VideoMAE, DisMo, and Meta’s V-JEPA 2 on motion-heavy evaluations while remaining weaker on appearance-dominated tasks such as HMDB51 and IARD. A motion-inversion probe shows the representation follows flipped or reversed motion while retaining the original answer on at most 1% of clips. The design also reduces temporal attention to 192 tokens and requires 456.1 GF encoder FLOPs, or 48% fewer than DisMo; experiments use NVIDIA hardware. Overall, the paper argues that motion sensitivity comes from the specific interaction between a narrow temporal bottleneck and first-frame reconstruction, not simply from adding temporal layers or diffusion objectives.

Original abstract

Comparisons in video self-supervised learning often evaluate complete training recipes rather than isolating the method itself: architecture, objective, data exposure, schedule, scale, and decoder capacity can all vary at once. This makes it hard to identify which choices yield motion-prioritized representations, whose gains concentrate on frame-to-frame change while retaining useful appearance. We address this with a matched $4 \times 6 = 24$ architecture-objective study at roughly 170M ~ 190M encoder scale on $\sim$1.7M OpenVid and Moments-in-Time v2 clips for 8 epochs, and propose TT-VidT. TT-VidT combines a DINOv3-initialized ViT-B/16 per-frame spatial path with a compact Temporal Transfer Layer, trained by Diff Compression to reconstruct target frames from a first-frame appearance anchor and frame-specific motion tokens. The sweep shows that TT3D with Diff Compression, not either component alone, enters the strongest motion-sensitive regime, and decoder ablations favor a compact video-pretrained decoder. In final comparison, TT-VidT leads Jester, Something-Something V2, ARID, and Diving48 fine-tuning simultaneously, improving over the strongest non-TT row by 54% ~ 121%, while using 48% fewer encoder FLOPs than DisMo and 55% fewer than VideoMAE or V-JEPA2. HMDB51, IARD, and EPIC-Kitchens bound the claim.

Read the original paper

More in Self-Supervised Learning

Browse all 22 papers →
01Self Supervised

Self-Play Pretraining with Zero Data

Aditya Cowsik, Kfir Dolev, Michael Y. Li, G. Bruno De Luca, Nourya Cohen, Noah D. Goodman, Yoav Levine

A learner and an RL-powered program generator teach each other from scratch, producing synthetic data that enables surprisingly meaningful transfer to natural datasets.

Read analysis
02Self Supervised

Strategically Diverse Sampling for Self-Training

Alexander Gurung, Esmeralda S. Whitammer, Mirella Lapata

Instead of training LLMs on many similar correct answers, this work shows that exposing them to diverse problem-solving strategies—even imperfect ones—can produce stronger models.

Read analysis