Vesta: A Generalist Embodied Reasoning Model
AuthorsJohan Bjorck, Zhiqi Li, Yunze Man, Jing Wang, An-Chieh Cheng, Sifei Liu, Shihao Wang, Zhiding Yu, Abhishek Badki, Stan Birchfield, Valts Blukis, Yevgen Chebotar, Siyi Chen, Sicong Leng, Yu-Cheng Chou, Tianli Ding, Boyi Li, Zhengyi Luo, Hang Su, Jonathan Tremblay, Tingwu Wang, Bowen Wen, Jimmy Wu, Xianghui Xie, Hanrong Ye, Hongxu Yin, K. R. Zentner, Liangyan Gui, Yu-Xiong Wang, Yuke Zhu, Linxi "Jim" Fan, Jan Kautz
Resources
Vesta is a robot brain that tries to do localization, navigation, spatial reasoning, and long-horizon planning in one generalist model instead of juggling multiple specialist systems.
Key results
Vesta average score on embodied cognition benchmarks
Vesta average score on localization benchmarks
Vesta average score on the offline real-robot action planning benchmark
Vesta success rate on R2R-CE val_unseen navigation
Vesta oracle success on R2R-CE val_unseen navigation
Average success improvement over actor-only baseline on real robotic tasks
What the paper found
NVIDIA’s Vesta is a unified embodied reasoning model that collapses localization, navigation, embodied question answering, and long-horizon action planning into a single Qwen3-VL-8B-based planner, instead of a brittle multi-model stack. The core technical idea is a curated supervised fine-tuning mixture biased toward spatial intelligence, navigation, grounding, embodied reasoning, and real-robot data, combined with a minimalist multimodal memory harness that interleaves retained image frames with a running text cache of prior subtasks. Across embodied cognition and localization benchmarks, Vesta reaches 68.7 average cognition and 69.9 average localization, outperforming RynnBrain and RoboBrain 2.5 in most categories; on offline real-robot planning it scores 75.4 average versus 38.5 for RoboBrain-2.5-8B. On navigation, it matches the specialist InternVLA-N1 with 55.5 SR and 61.4 OS on R2R-CE val_unseen. Most notably, on a bimanual YAM gripper robot with Gr00t-N1.6 as the low-level actor, Vesta improves success by 38.3% over actor-only control and 25% over a Qwen3-VL planner, showing that a single generalist planner can outperform specialized systems while remaining directly deployable in hierarchical robotic control.
Original abstract
Robots operating in open-world environments must seamlessly integrate localization, spatial reasoning, navigation, and long-horizon planning. While specialist models excel at individual tasks, deploying a multi-model stack is computationally expensive and prone to cascading errors. We present Vesta, a unified embodied generalist that consolidates these capabilities into a single foundation model. Our approach combines a diverse and massive curated corpus designed to induce spatial grounding and a simple multimodal memory harness that enables reasoning over extended time horizons. Across diverse benchmarks, Vesta on average beats individual SOTA baselines by >$20\%$ and beats an ensemble of per-category-best baselines by $>10\%$ -- thus demonstrating that a generalist model can match or exceed specialists. On real-world robotic tasks requiring memory and reasoning, Vesta improves task success by >35\%. Our work thus demonstrates that a single generalist is a feasible, scalable, and arguably preferable alternative to combining specialists.
Read the original paperMore in Embodied AI
Browse all 48 papers →GroundingPI: A Grounding Foundation Model towards Physical Intelligence with Visual Primitives
Qize Yu, Lianrui Fan, Boyu Chen, Jiaqi Liang, Xini Ding, Yue Chen, Zetian Song, Yuran Wang, Yi Zou, Kaixuan Wang, Tianxing Chen, Wenxuan Song, Bohan Zhou, Mingleyang Li, Siqiao Huang, Yuqi Ye, Caigao Jiang, Wei Wei, Ruihai Wu, Hang Zhang, Yixiao Ge, Shuchang Zhou, Shilong Liu, Xianming Liu, Ping Luo, Shiyu Huang
GroundingPI argues that fast, precise visual grounding should be the perceptual foundation for capable robots and autonomous vehicles.
MM-ABC: Towards Generalist Mobile Manipulation via Seeing, Coordinating and Imagining
Qiwei Liang, Guangyu Chen, Shaolong Zhu, Zikuan Xiao, Jinxuan Lu, Yifan Xie, Renjing Xu, Wenbo Ding, Tianxing Chen
MM-ABC is a generalist robot foundation model that helps mobile manipulators see their surroundings, coordinate arm and base motion, and imagine future actions for better performance.
Morphometric Imitation: From Morphology and Contact Aware Hand Retargeting to Sim-to-Real Visuomotor Policy
Tara Sadjadpour, Siming He, C. K. Wolfe, Haozhi Qi, Lea Wilken, S. Shankar Sastry, Claire Tomlin, Jitendra Malik
A three-stage system converts human hand demonstrations into robust, zero-shot real-robot dexterous manipulation policies across different hand morphologies.