NTH

PointZero: 3D Point Track Completion for Learning Transferable 3D Dynamics

AuthorsBardienus P. Duisterhof, Kaifeng Zhang, Adam Hung, Bowen Wen, Stan Birchfield, Yunzhu Li, Deva Ramanan, Jeffrey Ichnowski

AffiliationsCMU · Columbia · NVIDIA

September 24, 2026 2 min read
Watch on YouTube
The one-line take

PointZero learns general 3D motion dynamics from sparse point tracks rather than robot actions, then transfers that knowledge to manipulation and imitation tasks.

Key results

2.9M
Synthetic training dataset

Synthetic frames spanning deformable, articulated, and rigid interactions.

14
Real-world evaluation objects

Unseen objects used for zero-shot sim-to-real evaluation.

124
Real-world evaluation interactions

Human-object interactions collected across the evaluation objects.

26%
PGND zero-shot MDE reduction

Approximate reduction versus the strongest baseline across six PGND scenes.

88.2%
Imitation-learning average success

Average success with pre-training and auxiliary point-track supervision, compared with 80.5% from scratch.

6/7
Manipulation tasks at highest or joint-highest success

Tasks where PointZero matches or exceeds the baselines.

What the paper found

PointZero introduces 3D point-track completion as a robot-free pre-training objective for transferable 3D dynamics. Given one RGB-D observation and just 1–3 sparse point trajectories, it predicts dense future 3D tracks for every observed point across rigid, articulated, and deformable objects, without robot-action labels. Its architecture combines a Perceiver-IO encoder, DINOv2 visual features, and an alternating self- and cross-attention diffusion transformer, trained with flow matching or JiT-style x-prediction. The released synthetic dataset contains 2.9M frames, while a real-world evaluation covers 14 objects and 124 interactions, using NVIDIA FleX for deformable simulation and CoTracker3 plus FoundationStereo for tracking and depth. Zero-shot evaluation beats adapted baselines on 11 of 12 real-world metrics, and on the PGND benchmark PointZero-FM reduces average mean distance error by approximately 26% relative to the strongest baseline. After fine-tuning for imitation learning with only 20 labeled demonstrations per task, auxiliary point-track supervision raises average simulation success from 80.5% to 88.2%, while PointZero achieves the highest or joint-highest success on 6/7 manipulation tasks. The results position dense 3D trajectory prediction as a bridge between scalable, actionless video-style pre-training and robot-conditioned world models.

Original abstract

World models endow perceptual systems with the ability to predict how scenes evolve under interaction. They are most beneficial when trained on diverse volumes of data, to instill a rich prior into downstream applications. Existing methods typically require robot action labels to learn action-conditioned 3D dynamics, which excludes web video data from the training pool. We study 3D point track completion as a pre-training objective for learning transferable 3D dynamics without robot data. Given a single RGB-D observation and sparse partial 3D trajectories (tracks), we predict future 3D tracks of all observed points. We show this objective produces a rich 3D dynamics prior, without requiring robot action labels. We contribute a diverse dataset of 2.9 million synthetic frames spanning deformable, articulated, and rigid objects, and use it to train PointZero. We show that a flexible and expressive transformer, PointZero, outperforms prior methods on the same data. We demonstrate the utility of our pre-training objective by post-training PointZero for two downstream applications: (1) action-conditioned 3D dynamics prediction and (2) imitation learning. When fine-tuned to condition on end-effector pose, PointZero outperforms the baselines on the recent PGND 3D dynamics benchmark. When fine-tuned to predict robot actions and 3D tracks, PointZero outperforms or matches the baselines on 6/7 simulated and real-world robot manipulation tasks. We furthermore evaluate training PointZero from scratch to isolate the benefits of our proposed architecture from those of our proposed pre-training objective and dataset. We release the dataset, checkpoints, and full training recipe.

Read the original paper

More in World Models

Browse all 41 papers →
01World Model

4Director: Controlling Video World Models with Rigid 3D Geometry

Wei Cao, Hao Zhang, Vikram Voleti, Yuqun Wu, Mallikarjun B R, Shimon Vainer, Mark Boss, Yaoyao Liu

4Director makes video world models controllable by moving explicit 3D meshes through time while preserving realistic, consistent appearances.

Read analysis
02World Model

EMPIRIC: Experiment-Driven Learning of Residual World Models for Robot Planning

Yichao Liang, Amber Li, Dat Nguyen, Emily Bunnapradist, Michelangelo Naim, Sreela Kodali, Matteo Merler, Bowen Li, Kiran Gopinathan, Yiyun Liu, Nikhil Pimpalkhare, Joshua B. Tenenbaum, Adrian Weller, Zenna Tavares, Tom Silver, Kevin Ellis

EMPIRIC lets robots discover missing physics through targeted experiments and use the resulting interpretable world models to plan better.

Read analysis
03World Model

RoboCoach: World Models as Active Coaches for Compositional Robot Skills

Jiajun Liu, Yifan Chen, Yichao Liu, Jiayi Zhang, Ruoqu Chen, Shaoxuan Xie, Guocai Yao, Mengdi Xu, Sen Cui, Changshui Zhang

RoboCoach uses imagined robot failures to decide what demonstrations to request next, making long-horizon manipulation skills improve more efficiently.

Read analysis