NTH

What, Where, and How: Probing Spatiotemporal Representations in Video Foundation Models

AuthorsSharon S. Musa, Fereshteh Forghani, Harrish Thasarathan, Sonia Joseph, Matthew Kowal, Konstantinos G. Derpanis

September 8, 2026 2 min read
Watch on YouTube
The one-line take

This study maps where video models learn motion and anomalies, finds weak intuitive physics, and uses latent geometry to create smoother camera-motion trajectories.

Key results

93.8%
VideoMAE-v2 camera-motion ROC-AUC

Mean test ROC-AUC across 15 CameraBench camera-motion tasks.

91.0%
V-JEPA 2 Large camera-motion ROC-AUC

Mean test ROC-AUC across CameraBench camera-motion tasks.

90.2%
V-JEPA 2 Giant camera-motion ROC-AUC

Mean test ROC-AUC across CameraBench camera-motion tasks.

68.4%
V-JEPA 2 Giant anomaly ROC-AUC

Test ROC-AUC on UBnormal anomaly detection.

62.1%
VideoMAE-v2 Giant anomaly ROC-AUC

Test ROC-AUC on UBnormal anomaly detection.

What the paper found

This study examines what, where, and how temporal information is encoded in the self-supervised video foundation models V-JEPA 2, developed by Meta, and VideoMAE-v2. Using frozen representations and lightweight linear SVM probes across transformer depth, the researchers test CameraBench, IntPhys 2, and UBnormal for camera motion, intuitive physics, and anomaly detection. Camera motion is strongly represented: VideoMAE-v2 Giant reaches 93.8% mean test ROC-AUC, while V-JEPA 2 Large and Giant reach 91.0% and 90.2%, showing that scaling does not guarantee better generalization. Intuitive-physics performance remains near 50% ROC-AUC across layers, indicating little linearly accessible knowledge of solidity, permanence, continuity, or immutability in the more challenging IntPhys 2 benchmark. Anomaly detection is intermediate, with V-JEPA 2 Giant reaching 68.4% test ROC-AUC versus 62.1% for VideoMAE-v2 Giant, and its strongest signals emerging in deeper layers. Beyond classification, PCA of per-temporal-patch features from RealEstate10K reveals smooth, curved, low-dimensional trajectories that track camera motion, although some clips exhibit oscillatory geometry. The paper then introduces geometry-aware cubic-spline, or tangent-based, steering, which follows these empirical trajectories instead of cutting through latent space with linear interpolation. Nearest-neighbor retrieval across 200 evaluation clips shows spline steering consistently stays closer to the original representation manifold and produces smoother, more temporally coherent frame progressions, suggesting that video latent spaces are not only decodable but also geometrically navigable.

Original abstract

Self-supervised video foundation models learn rich spatiotemporal representations, yet it remains unclear what visual concepts these representations encode, where they emerge across transformer layers, and how they are geometrically organized. In this work, we tackle these three questions through a systematic layer-wise analysis of V-JEPA 2 and VideoMAE-v2. We leverage lightweight probes trained to discover three temporally grounded properties: (i) camera motion understanding, (ii) intuitive physics, and (iii) anomaly detection. Both models encode camera motion, with best results ($>90$ ROC AUC) emerging at 60-70% of network depth, and achieve moderate anomaly detection performance ($>60$ ROC AUC), but remain near chance on intuitive-physics tasks, suggesting a limited encoding of deeper physical reasoning. Beyond classification, we find that temporal features from individual videos form smooth low-dimensional trajectories in representation space, suggesting that camera motion is not only linearly decodable but also geometrically organized. Based on these results, we apply geometry-aware spline-based steering in the model's latent representations to interpolate camera motion, yielding steered videos with smoother trajectories and more coherent temporal progression than linear interpolation.

Read the original paper

More in Foundation Models

Browse all 47 papers →
01Foundation Model

How Much Is an AI Token Worth? Scaling Laws for Wild AI-Generated Web Text

Jenna Russell, Ben Glickenhaus, Katherine Thai, John Wieting, Mohit Iyyer, Max Spero, Bradley Emi

AI-generated web text can help language models at first, but beyond a tipping point it degrades performance on human writing, making data filtering and separate evaluation increasingly important.

Read analysis
02Foundation Model

TabFM: A Zero-Shot Foundation Model for Tabular Data

Weihao Kong, Erez Louidor Ilan, Shuxin Nie, Taman Narayan, Rajat Sen, Yichen Zhou, Deqing Fu, Samet Oymak, Abhimanyu Das

TabFM is a large synthetic-data-trained model that aims to make accurate tabular predictions instantly, without retraining for each new dataset.

Read analysis
03Foundation Model

When Do Biological Reasoning Models Use Their Biological Inputs?

Ada Fang, Nikitha Thoduguli, Lukas Fesser, Hanlin Zhang, Sham M. Kakade, Marinka Zitnik

The study finds that many biological reasoning systems appear to succeed without meaningfully using the biological inputs they were designed to reason over.

Read analysis