NTH

GroundingPI: A Grounding Foundation Model towards Physical Intelligence with Visual Primitives

AuthorsQize Yu, Lianrui Fan, Boyu Chen, Jiaqi Liang, Xini Ding, Yue Chen, Zetian Song, Yuran Wang, Yi Zou, Kaixuan Wang, Tianxing Chen, Wenxuan Song, Bohan Zhou, Mingleyang Li, Siqiao Huang, Yuqi Ye, Caigao Jiang, Wei Wei, Ruihai Wu, Hang Zhang, Yixiao Ge, Shuchang Zhou, Shilong Liu, Xianming Liu, Ping Luo, Shiyu Huang

Affiliations[

October 4, 2026 3 min read
Watch on YouTube
The one-line take

GroundingPI argues that fast, precise visual grounding should be the perceptual foundation for capable robots and autonomous vehicles.

Key results

4B
Model size

GroundingPI parameter scale

73.68%
Average grounding score

Mean across 34 grounding benchmarks

71.54%
GPT-6 Astra average

Comparison average across the grounding evaluation

24.8%
Maximum RoboTwin OOD improvement

Relative improvement over the strongest backbone

0.296 m
nuScenes open-loop L2 error

Average trajectory error using GroundingPI as the visual backbone

28.75%
RoboCasa-GR1 reduced-data success rate

Success rate using 50% of demonstrations

What the paper found

GroundingPI is a 4B-parameter grounding foundation model designed to give physical-intelligence systems precise visual primitives before they learn actions. It combines a MoonViT-V2 visual encoder with a Qwen3-4B decoder and represents points and bounding boxes as quantized coordinate tokens in a shared vocabulary, covering detection, referring expressions, pointing, OCR, GUI interaction, document layout, and visual prompting. Its staged training uses multimodal and spatial pretraining, supervised fine-tuning, and GRPO reinforcement learning with task-specific localization and text-geometry rewards. Across 34 grounding benchmarks and 44 baselines, GroundingPI averages 73.68%, exceeding the larger GPT-6 Astra at 71.54%; it also outperforms general-purpose backbones such as Qwen3-VL-4B and embodied systems associated with NVIDIA, including GR00T N1, in the controlled manipulation comparisons. As a downstream visual backbone, it improves robotic and driving performance: on RoboTwin 2.0, the strongest out-of-distribution gain reaches 24.8% relative, while on nuScenes its average open-loop trajectory error is 0.296 m. GroundingPI also improves action-data efficiency on RoboCasa-GR1, reaching 28.75% success with 50% of demonstrations, above every compared baseline trained with 75%. Ablations show that dense grounding supplies the main transferable spatial signal, while OCR may amplify fine-grained perceptual learning. The paper’s central claim is that perception-native foundations can complement broad reasoning models such as GPT-6 Astra by supporting faster, more reliable System 1 perception-action execution.

Original abstract

Precise grounding matters. It specifies which object is the target and where that object is, even in clutter and for tiny objects, and it has to be fast enough for closed-loop control. Yet vision-language-action (VLA) and world-action models (WAMs) take perception from general-purpose vision-language and video-generation backbones, which still fail in these settings. We introduce GroundingPI, a 4B grounding foundation model that generates points and boxes as quantized coordinates in a shared vocabulary. Training combines multimodal and spatial pretraining, supervised fine-tuning, and reinforcement learning with GRPO, using supervision from public datasets and dedicated data engines. Against 44 baselines across 34 grounding benchmarks spanning 11 perceptual capabilities, GroundingPI establishes a new state of the art, averaging 73.68%, above the larger GPT-6 Astra (71.54%). As a downstream visual backbone, GroundingPI improves performance on robotic manipulation and autonomous driving. On RoboTwin 2.0, it outperforms every mainstream backbone we evaluate in all four out-of-distribution settings, by up to 24.8% relative to the strongest backbone. On RoboCasa-GR1, GroundingPI trained with 50% of the demonstrations outperforms those baselines trained with 75%. On nuScenes, used as the visual backbone, GroundingPI attains an average open-loop L2 error of 0.296 m. We systematically analyze GroundingPI's pretraining in scale and data composition. Downstream autonomous driving and robotic manipulation improve as the pretraining is scaled. Analyzing the data recipe across these 11 perceptual capabilities shows dense grounding's substantial benefits for both, and OCR's potential as a catalyst for perceptual learning. These results support grounding as a perceptual foundation, and dedicated perceptual pretraining as a promising direction for foundation models of physical intelligence.

Read the original paper

More in Embodied AI

Browse all 48 papers →
03Embodied Ai

Agent as Policy for Robotic Manipulation

Mengzhao Jia, Yang Lin, Xixin Zhang, Zhihan Zhang, Xiaobai Liu, Meng Jiang

A general-purpose agent becomes a robot policy by writing and adapting its own programs while interacting with the physical world.

Read analysis