NTH

Track2View: 4D-Consistent Camera-Controlled Video Generation via Paired 3D Point Tracks

AuthorsFeng Qiao, Zhaochong An, Zhexiao Xiong, Serge Belongie, Nathan Jacobs

June 18, 2026 2 min read
Watch on YouTube
The one-line take

Track2View lets a video diffusion model re-render scenes from new camera paths by using paired 3D point tracks to keep motion and appearance consistent across time.

Key results

400
benchmark videos

RealCam-Vid evaluation set size

200
dynamic scenes

MiraData subset in the benchmark

26.82
FID

Track2View score at 81 frames on RealCam-Vid

93.22
CLIP-V

Track2View view-synchronization score at 81 frames

0.695K
Mat.Pix.

Track2View matched keypoints at 81 frames

161
query frames in paired-track extraction

Temporally concatenated sequence length used with SpatialTrackerV2

What the paper found

Track2View is a camera-controlled video re-generation framework that targets 4D consistency by conditioning a video diffusion transformer on paired 3D point tracks instead of per-frame camera embeddings or noisy rendered geometry. Built on the WAN-2.1 video diffusion transformer, it uses a dual-view track conditioner with parameter-free bilinear sampling and scattering, plus an 8-layer temporal aggregation transformer, to propagate sparse track information across source and target views while preserving explicit spatial and temporal correspondences. The training pipeline extracts one-to-one paired tracks by running SpatialTrackerV2 on a 161-frame temporally concatenated sequence from the MultiCamVideo synthetic dataset, then fine-tunes the model with LoRA rank 64 for 216K optimization steps. On the 400-video RealCam-Vid benchmark spanning 200 static RealEstate10K scenes and 200 dynamic MiraData scenes, Track2View sets the best overall camera control and view-synchronization results, reaching FID 26.82, CLIP-V 93.22, and Mat.Pix. 0.695K at 81 frames, while reducing rotation error by 30–65% and translation error by 61–72% versus leading baselines. It also scores best or second-best across all six VBench dimensions, indicating that stronger geometric control does not sacrifice perceptual video quality.

Original abstract

Re-rendering an existing video from a novel camera viewpoint requires the output to follow the prescribed camera trajectory while preserving the appearance and dynamics of the original scene across every frame. Existing methods rely on per-frame pose embeddings, noisy point-cloud renderings, or implicit learned correspondences, none of which provides an explicit, temporally continuous link between source and target pixels. We propose Track2View, which conditions a video diffusion transformer on paired 3D point tracks: sparse trajectories of scene points projected into both the source and target camera views. These tracks provide explicit spatiotemporal correspondences that are temporally continuous by construction, encoding what content should appear where and when. At the core of Track2View is a dual-view track conditioner that transfers visual context from source to target view through parameter-free geometric operations and learned temporal aggregation, ensuring generalization to arbitrary camera trajectories without memorizing specific motions. We further introduce a data curation pipeline that extracts one-to-one track correspondences by running a 3D point tracker on temporally concatenated multi-camera view pairs. On a 400-video benchmark spanning static and dynamic scenes, Track2View achieves state-of-the-art results across visual quality, view synchronization, and camera accuracy, reducing rotation error by 30-65% and translation error by 61-72% relative to leading baselines. Project page is available at this https URL: https://qjizhi.github.io/track2view

Read the original paper

More in Computer Vision

Browse all 58 papers →
02Cv

DyRAD: Radar Novel View Synthesis for Dynamic Driving Scenes

Merav Keidar, Tomer Borreda, Rajalakshmi Nandakumar, Or Litany

DyRAD builds moving radar views of driving scenes by combining tracked object motion with the radar’s physics, enabling more realistic and transferable autonomy testing.

Read analysis
03Cv

OmniTaskonomy: When Does Visual Generation Improve Visual Understanding?

Jiaxin Ge, Yiming Qin, Ji Xie, Haozhe Jiang, Xiaochuang Han, Junyi Zhang, Andrew Dai, Yinfei Yang, Jitendra Malik, Ranjay Krishna, Sewon Min, Haiwen Feng, Le Xue, Baifeng Shi, Trevor Darrell, XuDong Wang

This work maps when training models to generate images can make them better at understanding images, revealing both intuitive and surprising task-to-task benefits.

Read analysis