3D HAMSTER: Bridging Planning and Control in Hierarchical Vision Language Action Models through 3D Trajectory Guidance
AuthorsDongyoon Hwang, Byungkun Lee, Dongjin Kim, Hyojin Jang, Hoiyeong Jin, Jueun Mun, Minho Park, Hojoon Lee, Hyunseung Kim, Jaegul Choo
Resources
3D HAMSTER makes robot planners output depth-aware 3D waypoints instead of flat 2D paths, helping language-guided manipulation work better in real-world and visually shifting environments.
Key results
Evaluation set used for 3D trajectory prediction
3D HAMSTER on DroidSpatial-Bench
Strong open-source baseline on DroidSpatial-Bench
3D HAMSTER across 11 simulated tasks
2D-guided baseline in Colosseum
3D HAMSTER on Franka Panda
What the paper found
3D HAMSTER, from the KAIST and KRAFTON AI collaboration, addresses a core failure mode in hierarchical vision-language-action systems: planners output 2D waypoints that low-level 3D policies must unproject, creating depth ambiguity and the “graffiti effect” on point clouds. The method upgrades a pretrained Qwen3-VL-8B-Instruct planner with a dedicated depth encoder and a dense depth reconstruction objective, so it predicts metrically reliable (u, v, d) trajectories that are directly consumable by a pointcloud-based 3D policy, specifically 3D FlowMatch Actor. On DroidSpatial-Bench, built from 148 held-out DROID pick-and-place episodes, the full model reaches 65.5% on the strict 10 cm “Both” metric, versus 60.1% for RoboBrain-2.5-8B and 29.7% for Gemini-3.0-Pro at 5 cm. In Colosseum simulation, 3D guidance raises average success to 44.8%, outperforming 2D HAMSTER at 38.8% and the unguided baseline at 36.6%, with especially strong gains under lighting, object, and texture perturbations. On a Franka Panda robot, the 3D planner achieves 80% on button pressing, 68% on pouring, and 62% on pick-and-place, beating both HAMSTER and the monolithic pi0.5 model on average, and showing the largest advantage on pouring, where precise depth control matters most.
Original abstract
Hierarchical Vision-Language-Action (VLA) models decouple high-level planning from low-level control to improve generalization in robot manipulation. Recent work in this paradigm uses 2D end-effector trajectories predicted by a Vision-Language Model (VLM) as explicit guidance for a downstream policy. However, state-of-the-art low-level policies operate in 3D metric space on point clouds, and feeding them 2D guidance that lacks depth forces each waypoint to be assigned the depth of whatever scene surface lies beneath it, producing geometrically distorted trajectories. We propose 3D HAMSTER, a hierarchical framework that closes this gap by having the planner directly output metrically reliable 3D trajectories. We augment a VLM with a dedicated depth encoder and a dense depth reconstruction objective to predict 3D waypoint sequences, which are directly integrated into a pointcloudbased low-level policy. Across 3D trajectory prediction, simulation, and real-world manipulation, 3D HAMSTER consistently outperforms proprietary VLMs and 2D-guided baselines, with the largest gains under appearance-altering shifts and unseen language, spatial, and visual conditions. The project page is available at https://davian-robotics.github.io/3D_HAMSTER/.
Read the original paperMore in Robotics
Browse all 50 papers →JAMB: Joint Action-Motion Diffusion for Bimanual Manipulation
Chuyang Xiao, Peilin Meng, David Held
JAMB helps two robot arms coordinate by jointly imagining their future movements and the changing 3D scene before acting.
Rolling-WAM: World Action Models with Rolling Imagination
Yinghua Zhou, Junjie Ye, Yiqi Zhao, Hao Dong, Celina Shiyu Wang, Ruohai Ge, Tingyi Yang, Basile Van Hoorick, Gaurav Sukhatme, Vitor Guizilini, Yue Wang
Rolling-WAM keeps future robot actions partially imagined and refined over time, making world-model-based manipulation replan 4.5 times faster.
Training-free Behavior Cloning
Maximilian Adang, Timothy Chen, Lars Osterberg, Aiden Swann, Mac Schwager
A fast, training-free robot controller reuses and corrects demonstration trajectories to deliver traceable behavior at real-time speeds.