Go-with-the-Track: Video Compositing and Motion Control with Point Tracking
AuthorsKoichi Namekata, Yash Kant, Zhizheng Liu, Ryan D Burgert, Yuancheng Xu, Kuan Heng Lin, Emmett Steven, Julien Philip, Li Ma, Andrea Vedaldi, Paul Debevec, Ning Yu
Resources
This paper lets creators control video generation by tracking points across reference images and frames, enabling more precise compositing and camera motion in a single model.
Key results
training corpus after automated filtering
training mix across synthetic and real video data
target video frames per training sample
training resolution for the video model
conditioning scale supported by the point-track design
best reported visual fidelity on DAVIS dense tracks
What the paper found
Go-with-the-Track, from Netflix, Eyeline Labs, and the University of Oxford, unifies reference-image compositing and motion control in a single video diffusion transformer by conditioning on multiple reference images plus reference-anchored point-tracks. The core idea is a spatially-aware point-track embedding: each trajectory is encoded from its full coordinate sequence with a coordinate-wise MLP and temporal max pooling, so embedding similarity reflects spatial proximity instead of using random IDs. A lightweight adapter then maps up to 15,000 pixel-space tracks into Wan 2.1/2.2’s 4×16×16 patchified latent space without the motion detail loss caused by naive subsampling. Training mixes real video, static-scene, and synthetic data, including approximately 500K curated videos and a 3:7 synthetic-to-real ratio, with 49-frame clips at 480×832 and up to 24K iterations. On DAVIS 2017 and TAPVid3D-ADT, the method consistently beats baselines such as DiffusionAsShader, Go-with-the-Flow, Tora, ATI, and Wan-Move; for example, on DAVIS dense tracks it reaches FID 28.00, FVD 322.8, and EPE 7.709, while on TAPVid3D-ADT dense tracks it reaches FID 41.00, FVD 314.3, and EPE 4.429. In a 45-participant user study, it wins 46.2% for motion following, 43.5% for subject preservation, and 44.3% overall quality, and it enables mesh compositing, keypoint-driven subject transfer, and camera retargeting for both static and dynamic scenes.
Original abstract
Filmmaking demands precise motion control and reference image compositing -- capabilities that existing methods treat separately. Point-track-conditioned image-to-video models restrict content insertion to the first frame, while reference-to-video models lack fine-grained spatial-temporal control over how reference content integrates across frames. We present Go-with-the-Track, which unifies both capabilities by jointly conditioning on multiple reference images and reference-anchored point-tracks -- extending conventional point-tracks to explicitly establish correspondences between generated frames and reference images, thus enabling precise compositing and motion control throughout the video. To achieve this, we introduce spatially-aware point-track embeddings that encode the full sequence of point-track coordinates using a coordinate-wise MLP followed by temporal pooling. This representation captures the spatial characteristics of each point-track (serving as a unique identifier), while the embedding similarity correlates directly with spatial proximity, enhancing the model's ability to distinguish and associate point-tracks. We inject these point-track embeddings into a video diffusion transformer via a lightweight adapter, resolving the pixel-to-patch resolution mismatch while avoiding the substantial motion detail loss inherent in naive point-track subsampling. We use a hybrid training strategy to train jointly on dynamic, static, and synthetic scene video datasets to boost motion controllability. Experiments demonstrate that Go-with-the-Track achieves superior motion and reference control in a single model and enables new capabilities: multi-reference conditioned video generation with point-track driven compositing, as well as camera control for both static and dynamic scenes. Project Page: https://eyeline-labs.github.io/Go-with-the-Track/
Read the original paperMore in Computer Vision
Browse all 58 papers →All-in-One Multilingual Scene Text Recognition with Script-aware Mixture-of-Experts
Xingsong Ye, Yongkun Du, Jiaxin Zhang, Zhixian Li, Chong Sun, Chen Li, Jing Lyu, Lianwen Jin, Zhineng Chen
A lightweight script-aware mixture-of-experts model brings more accurate, scalable multilingual scene text recognition to many languages and scripts.
DyRAD: Radar Novel View Synthesis for Dynamic Driving Scenes
Merav Keidar, Tomer Borreda, Rajalakshmi Nandakumar, Or Litany
DyRAD builds moving radar views of driving scenes by combining tracked object motion with the radar’s physics, enabling more realistic and transferable autonomy testing.
OmniTaskonomy: When Does Visual Generation Improve Visual Understanding?
Jiaxin Ge, Yiming Qin, Ji Xie, Haozhe Jiang, Xiaochuang Han, Junyi Zhang, Andrew Dai, Yinfei Yang, Jitendra Malik, Ranjay Krishna, Sewon Min, Haiwen Feng, Le Xue, Baifeng Shi, Trevor Darrell, XuDong Wang
This work maps when training models to generate images can make them better at understanding images, revealing both intuitive and surprising task-to-task benefits.