NTH

3DWay: Generalizing Robot Manipulation via 3D Consistent Waypoints

AuthorsZiqin Huang, Yingyue Li, Chenyangguang Zhang, Ruida Zhang, Yuxin Chen, Gu Wang, Xingyu Liu, Masayoshi Tomizuka, Xiangyang Ji

September 13, 2026 2 min read
Watch on YouTube
The one-line take

3DWay helps robots plan more reliably by turning multi-view visual predictions into geometrically consistent 3D movement waypoints.

Key results

134k
Training trajectories

Combined processed trajectories from RLBench, DROID, and RH20T.

15B
NVILA backbone

Default VLM used to generate multi-view-consistent 2D waypoints.

64.0%
Unseen RLBench success

Success rate achieved by direct 3DWay-TD execution on unseen tasks.

6.4%
VLABench improvement

Average success-rate advantage over π0.5.

65.8%
Real-world basic-task success

Average success rate for waypoint-augmented π0, compared with 21.7% for standard fine-tuning.

What the paper found

3DWay addresses a central weakness of robot manipulation systems: 2D trajectories leave depth and free-space positions ambiguous, while depth-back-projected 2.5D paths remain sensitive to noise. Its solution is to fine-tune a pretrained vision-language model, using NVILA variants, to predict multi-view-consistent 2D waypoints from RGB images and language, then recover deterministic 3D tool-center-point trajectories through calibrated geometric triangulation. Training combines RLBench, DROID, and RH20T, covering around 290 tasks and 134k trajectories, with Ramer-Douglas-Peucker simplification and pixel-coordinate supervision. The resulting 3DWay representation can drive simple tasks directly or guide foundation vision-language-action models through an adaptive local waypoint window; integration with π0 preserves its learned action patterns while adding explicit spatial grounding. Using NVILA-15B, direct 3DWay-TD execution reached 64.0% success on unseen RLBench tasks and exceeded π0.5 by 6.4% on VLABench’s average score. In real-world experiments, waypoint-augmented π0 improved average basic-task success from 21.7% to 65.8%, while also handling novel objects and abstract instructions. Ablations show that multi-view consistency is essential, although the current system mainly models translation, assumes calibrated dual cameras, and does not yet represent end-effector orientation or dynamic adaptation.

Original abstract

Intermediate representations are key to bridging the modality gap between generalizable manipulation policies and large-scale pretrained vision-language models (VLMs). Among these, trajectory-based representations compactly represent motion-relevant cues, yet most existing approaches predict trajectories in 2D image space, resulting in intrinsic 3D ambiguity. Moreover, using 2D trajectories with depth still leaves the free-space waypoints ambiguous, limiting reliable 3D reasoning. To address this, we propose predicting 3D consistent waypoints (3DWay) from multi-view images. By reformulating 3D waypoints prediction as generating multi-view consistent 2D waypoints followed by geometric triangulation, we enable explicit 3D motion specification while preserving the strong priors of pretrained VLMs. The predicted waypoints can guide existing VLA models for better generalization or be directly executed on simple tasks. Extensive experiments show that 3DWay substantially improves 3D spatial grounding and vision-language reasoning, demonstrating strong potential for generalizable robot manipulation. Codes will be released at https://github.com/ziqin-h/3DWay.

Read the original paper

More in Robotics

Browse all 50 papers →
02Robotics

Rolling-WAM: World Action Models with Rolling Imagination

Yinghua Zhou, Junjie Ye, Yiqi Zhao, Hao Dong, Celina Shiyu Wang, Ruohai Ge, Tingyi Yang, Basile Van Hoorick, Gaurav Sukhatme, Vitor Guizilini, Yue Wang

Rolling-WAM keeps future robot actions partially imagined and refined over time, making world-model-based manipulation replan 4.5 times faster.

Read analysis
03Robotics

Training-free Behavior Cloning

Maximilian Adang, Timothy Chen, Lars Osterberg, Aiden Swann, Mac Schwager

A fast, training-free robot controller reuses and corrects demonstration trajectories to deliver traceable behavior at real-time speeds.

Read analysis