GroundingPI: A Grounding Foundation Model towards Physical Intelligence with Visual Primitives
AuthorsQize Yu, Lianrui Fan, Boyu Chen, Jiaqi Liang, Xini Ding, Yue Chen, Zetian Song, Yuran Wang, Yi Zou, Kaixuan Wang, Tianxing Chen, Wenxuan Song, Bohan Zhou, Mingleyang Li, Siqiao Huang, Yuqi Ye, Caigao Jiang, Wei Wei, Ruihai Wu, Hang Zhang, Yixiao Ge, Shuchang Zhou, Shilong Liu, Xianming Liu, Ping Luo, Shiyu Huang
Affiliations[
Resources
GroundingPI argues that fast, precise visual grounding should be the perceptual foundation for capable robots and autonomous vehicles.
Key results
GroundingPI parameter scale
Mean across 34 grounding benchmarks
Comparison average across the grounding evaluation
Relative improvement over the strongest backbone
Average trajectory error using GroundingPI as the visual backbone
Success rate using 50% of demonstrations
What the paper found
GroundingPI is a 4B-parameter grounding foundation model designed to give physical-intelligence systems precise visual primitives before they learn actions. It combines a MoonViT-V2 visual encoder with a Qwen3-4B decoder and represents points and bounding boxes as quantized coordinate tokens in a shared vocabulary, covering detection, referring expressions, pointing, OCR, GUI interaction, document layout, and visual prompting. Its staged training uses multimodal and spatial pretraining, supervised fine-tuning, and GRPO reinforcement learning with task-specific localization and text-geometry rewards. Across 34 grounding benchmarks and 44 baselines, GroundingPI averages 73.68%, exceeding the larger GPT-6 Astra at 71.54%; it also outperforms general-purpose backbones such as Qwen3-VL-4B and embodied systems associated with NVIDIA, including GR00T N1, in the controlled manipulation comparisons. As a downstream visual backbone, it improves robotic and driving performance: on RoboTwin 2.0, the strongest out-of-distribution gain reaches 24.8% relative, while on nuScenes its average open-loop trajectory error is 0.296 m. GroundingPI also improves action-data efficiency on RoboCasa-GR1, reaching 28.75% success with 50% of demonstrations, above every compared baseline trained with 75%. Ablations show that dense grounding supplies the main transferable spatial signal, while OCR may amplify fine-grained perceptual learning. The paper’s central claim is that perception-native foundations can complement broad reasoning models such as GPT-6 Astra by supporting faster, more reliable System 1 perception-action execution.
Original abstract
Precise grounding matters. It specifies which object is the target and where that object is, even in clutter and for tiny objects, and it has to be fast enough for closed-loop control. Yet vision-language-action (VLA) and world-action models (WAMs) take perception from general-purpose vision-language and video-generation backbones, which still fail in these settings. We introduce GroundingPI, a 4B grounding foundation model that generates points and boxes as quantized coordinates in a shared vocabulary. Training combines multimodal and spatial pretraining, supervised fine-tuning, and reinforcement learning with GRPO, using supervision from public datasets and dedicated data engines. Against 44 baselines across 34 grounding benchmarks spanning 11 perceptual capabilities, GroundingPI establishes a new state of the art, averaging 73.68%, above the larger GPT-6 Astra (71.54%). As a downstream visual backbone, GroundingPI improves performance on robotic manipulation and autonomous driving. On RoboTwin 2.0, it outperforms every mainstream backbone we evaluate in all four out-of-distribution settings, by up to 24.8% relative to the strongest backbone. On RoboCasa-GR1, GroundingPI trained with 50% of the demonstrations outperforms those baselines trained with 75%. On nuScenes, used as the visual backbone, GroundingPI attains an average open-loop L2 error of 0.296 m. We systematically analyze GroundingPI's pretraining in scale and data composition. Downstream autonomous driving and robotic manipulation improve as the pretraining is scaled. Analyzing the data recipe across these 11 perceptual capabilities shows dense grounding's substantial benefits for both, and OCR's potential as a catalyst for perceptual learning. These results support grounding as a perceptual foundation, and dedicated perceptual pretraining as a promising direction for foundation models of physical intelligence.
Read the original paperMore in Embodied AI
Browse all 48 papers →MM-ABC: Towards Generalist Mobile Manipulation via Seeing, Coordinating and Imagining
Qiwei Liang, Guangyu Chen, Shaolong Zhu, Zikuan Xiao, Jinxuan Lu, Yifan Xie, Renjing Xu, Wenbo Ding, Tianxing Chen
MM-ABC is a generalist robot foundation model that helps mobile manipulators see their surroundings, coordinate arm and base motion, and imagine future actions for better performance.
Morphometric Imitation: From Morphology and Contact Aware Hand Retargeting to Sim-to-Real Visuomotor Policy
Tara Sadjadpour, Siming He, C. K. Wolfe, Haozhi Qi, Lea Wilken, S. Shankar Sastry, Claire Tomlin, Jitendra Malik
A three-stage system converts human hand demonstrations into robust, zero-shot real-robot dexterous manipulation policies across different hand morphologies.
Agent as Policy for Robotic Manipulation
Mengzhao Jia, Yang Lin, Xixin Zhang, Zhihan Zhang, Xiaobai Liu, Meng Jiang
A general-purpose agent becomes a robot policy by writing and adapting its own programs while interacting with the physical world.