NTH

Go-with-the-Track: Video Compositing and Motion Control with Point Tracking

AuthorsKoichi Namekata, Yash Kant, Zhizheng Liu, Ryan D Burgert, Yuancheng Xu, Kuan Heng Lin, Emmett Steven, Julien Philip, Li Ma, Andrea Vedaldi, Paul Debevec, Ning Yu

June 26, 2026 2 min read
Watch on YouTube
The one-line take

This paper lets creators control video generation by tracking points across reference images and frames, enabling more precise compositing and camera motion in a single model.

Key results

500K
curated videos

training corpus after automated filtering

3:7
synthetic-to-real ratio

training mix across synthetic and real video data

49
clip length

target video frames per training sample

480x832
input resolution

training resolution for the video model

15000
max point-tracks

conditioning scale supported by the point-track design

28.00
DAVIS dense FID

best reported visual fidelity on DAVIS dense tracks

What the paper found

Go-with-the-Track, from Netflix, Eyeline Labs, and the University of Oxford, unifies reference-image compositing and motion control in a single video diffusion transformer by conditioning on multiple reference images plus reference-anchored point-tracks. The core idea is a spatially-aware point-track embedding: each trajectory is encoded from its full coordinate sequence with a coordinate-wise MLP and temporal max pooling, so embedding similarity reflects spatial proximity instead of using random IDs. A lightweight adapter then maps up to 15,000 pixel-space tracks into Wan 2.1/2.2’s 4×16×16 patchified latent space without the motion detail loss caused by naive subsampling. Training mixes real video, static-scene, and synthetic data, including approximately 500K curated videos and a 3:7 synthetic-to-real ratio, with 49-frame clips at 480×832 and up to 24K iterations. On DAVIS 2017 and TAPVid3D-ADT, the method consistently beats baselines such as DiffusionAsShader, Go-with-the-Flow, Tora, ATI, and Wan-Move; for example, on DAVIS dense tracks it reaches FID 28.00, FVD 322.8, and EPE 7.709, while on TAPVid3D-ADT dense tracks it reaches FID 41.00, FVD 314.3, and EPE 4.429. In a 45-participant user study, it wins 46.2% for motion following, 43.5% for subject preservation, and 44.3% overall quality, and it enables mesh compositing, keypoint-driven subject transfer, and camera retargeting for both static and dynamic scenes.

Original abstract

Filmmaking demands precise motion control and reference image compositing -- capabilities that existing methods treat separately. Point-track-conditioned image-to-video models restrict content insertion to the first frame, while reference-to-video models lack fine-grained spatial-temporal control over how reference content integrates across frames. We present Go-with-the-Track, which unifies both capabilities by jointly conditioning on multiple reference images and reference-anchored point-tracks -- extending conventional point-tracks to explicitly establish correspondences between generated frames and reference images, thus enabling precise compositing and motion control throughout the video. To achieve this, we introduce spatially-aware point-track embeddings that encode the full sequence of point-track coordinates using a coordinate-wise MLP followed by temporal pooling. This representation captures the spatial characteristics of each point-track (serving as a unique identifier), while the embedding similarity correlates directly with spatial proximity, enhancing the model's ability to distinguish and associate point-tracks. We inject these point-track embeddings into a video diffusion transformer via a lightweight adapter, resolving the pixel-to-patch resolution mismatch while avoiding the substantial motion detail loss inherent in naive point-track subsampling. We use a hybrid training strategy to train jointly on dynamic, static, and synthetic scene video datasets to boost motion controllability. Experiments demonstrate that Go-with-the-Track achieves superior motion and reference control in a single model and enables new capabilities: multi-reference conditioned video generation with point-track driven compositing, as well as camera control for both static and dynamic scenes. Project Page: https://eyeline-labs.github.io/Go-with-the-Track/

Read the original paper

More in Computer Vision

Browse all 58 papers →
02Cv

DyRAD: Radar Novel View Synthesis for Dynamic Driving Scenes

Merav Keidar, Tomer Borreda, Rajalakshmi Nandakumar, Or Litany

DyRAD builds moving radar views of driving scenes by combining tracked object motion with the radar’s physics, enabling more realistic and transferable autonomy testing.

Read analysis
03Cv

OmniTaskonomy: When Does Visual Generation Improve Visual Understanding?

Jiaxin Ge, Yiming Qin, Ji Xie, Haozhe Jiang, Xiaochuang Han, Junyi Zhang, Andrew Dai, Yinfei Yang, Jitendra Malik, Ranjay Krishna, Sewon Min, Haiwen Feng, Le Xue, Baifeng Shi, Trevor Darrell, XuDong Wang

This work maps when training models to generate images can make them better at understanding images, revealing both intuitive and surprising task-to-task benefits.

Read analysis