TT-VidT: Decoupling the Temporal Axis for Efficient Motion-Centric Video Pretraining
AuthorsShih-Ying Yeh, Daniel Z. Kaplan, Xuehai Wang, Fu-En Yang, Min-Hung Chen, Shang-Hong Lai
AffiliationsNational Tsing Hua University Comfy Org Research realiz.ai · Karolinska Institutet Stockholm University NVIDIA†Corresponding author: kohaku@kblueleaf.netProject page: https://kohakublueleaf.github.io/TTVidT/
Resources
TT-VidT pretrains video models to focus on motion while preserving appearance, achieving strong action-recognition results with substantially lower compute.
Key results
Matched sweep covering four encoders and six objectives
Approximate OpenVid and Moments-in-Time v2 clips
Epochs under the shared training recipe
TT-VidT frozen attentive-probe score
TT-VidT frozen attentive-probe score
Encoder-only computation in GFLOPs
What the paper found
TT-VidT targets motion-centric video representations by separating appearance from temporal change during self-supervised pretraining. Its encoder combines a DINOv3-initialized ViT-B/16 spatial path with a compact Temporal Transfer Layer, while Diff Compression reconstructs each target frame from a wide first-frame appearance anchor and only 8 frame-specific motion tokens. The study evaluates 24 architecture-objective combinations using approximately 1.7M clips from OpenVid and Moments-in-Time v2 over 8 epochs, isolating the effect of architecture and objective under a matched recipe. The strongest pairing is TT3D with Diff Compression: unlike either component alone, it reaches 73.25 on Jester, 25.92 on Something-Something V2, 37.47 on ARID, and 18.63 after Diving48 fine-tuning, leading VideoMAE, DisMo, and Meta’s V-JEPA 2 on motion-heavy evaluations while remaining weaker on appearance-dominated tasks such as HMDB51 and IARD. A motion-inversion probe shows the representation follows flipped or reversed motion while retaining the original answer on at most 1% of clips. The design also reduces temporal attention to 192 tokens and requires 456.1 GF encoder FLOPs, or 48% fewer than DisMo; experiments use NVIDIA hardware. Overall, the paper argues that motion sensitivity comes from the specific interaction between a narrow temporal bottleneck and first-frame reconstruction, not simply from adding temporal layers or diffusion objectives.
Original abstract
Comparisons in video self-supervised learning often evaluate complete training recipes rather than isolating the method itself: architecture, objective, data exposure, schedule, scale, and decoder capacity can all vary at once. This makes it hard to identify which choices yield motion-prioritized representations, whose gains concentrate on frame-to-frame change while retaining useful appearance. We address this with a matched $4 \times 6 = 24$ architecture-objective study at roughly 170M ~ 190M encoder scale on $\sim$1.7M OpenVid and Moments-in-Time v2 clips for 8 epochs, and propose TT-VidT. TT-VidT combines a DINOv3-initialized ViT-B/16 per-frame spatial path with a compact Temporal Transfer Layer, trained by Diff Compression to reconstruct target frames from a first-frame appearance anchor and frame-specific motion tokens. The sweep shows that TT3D with Diff Compression, not either component alone, enters the strongest motion-sensitive regime, and decoder ablations favor a compact video-pretrained decoder. In final comparison, TT-VidT leads Jester, Something-Something V2, ARID, and Diving48 fine-tuning simultaneously, improving over the strongest non-TT row by 54% ~ 121%, while using 48% fewer encoder FLOPs than DisMo and 55% fewer than VideoMAE or V-JEPA2. HMDB51, IARD, and EPIC-Kitchens bound the claim.
Read the original paperMore in Self-Supervised Learning
Browse all 22 papers →Self-Play Pretraining with Zero Data
Aditya Cowsik, Kfir Dolev, Michael Y. Li, G. Bruno De Luca, Nourya Cohen, Noah D. Goodman, Yoav Levine
A learner and an RL-powered program generator teach each other from scratch, producing synthetic data that enables surprisingly meaningful transfer to natural datasets.
Strategically Diverse Sampling for Self-Training
Alexander Gurung, Esmeralda S. Whitammer, Mirella Lapata
Instead of training LLMs on many similar correct answers, this work shows that exposing them to diverse problem-solving strategies—even imperfect ones—can produce stronger models.
Human-JEPA: A Human-Centric Vision Model that Perceives and Anticipates
Hui Wei, Licai Sun, Guoying Zhao
Human-JEPA aims to help vision systems understand people now and predict what they will do next using one efficient self-supervised video model.