NTH

Adaptive Vision-Language Grasping via Composable Foundation Priors and Generalizable Grasp Synthesis

AuthorsSixu Yan, Shikang Wang, Binhua Huang, Xuanlai Tang, Guohua Fan, Fan Huang, Haoxuan Li, Yongkang Li, Yuhan Li, Bencheng Liao, Zeyu Zhang, Wenyu Liu, Hangxin Liu, Xinggang Wang

September 13, 2026 2 min read
Watch on YouTube
The one-line take

AdaRoboVLG lets robots adapt how they grasp objects by combining a reusable physical grasping policy with task-specific knowledge from foundation models.

Key results

4.4M
Simulated grasp trials

Annotated Isaac Sim trials used to train the hand-agnostic grasp decision model.

88.4%
DexGraspNet 2.0 clutter success

Average success rate across three robotic hands and three clutter levels.

94.0%
Functional-part accuracy

Open-set grasp functionality inference accuracy.

81.3%
GraspClutter6D language-guided success

Average success rate for language-guided target grasping.

83.3%
Real-world static-clutter success

Success across 510 trials involving 102 everyday objects.

89.7%
Dynamic-disturbance success

Average success rate under four human-induced motion and occlusion disturbances.

What the paper found

AdaRoboVLG addresses the brittleness of end-to-end vision-language grasping by separating task understanding from physical grasp synthesis. Its structured grasp interface combines object geometry, Contact Grasp Representations encoding contact point, approach and closing directions, grasp width, and taxonomy-compatible grasp types. A hand-agnostic base policy maps these primitives to executable poses for different robotic hands, then uses a PointTransformer-based hand-object interaction model and force-closure stability estimation to rank candidates. Three composable priors modify the same interface: a spatial prior uses DINOv3 features, multi-view RGB-D, and Contact-GraspNet to generate clutter-aware contact constraints; a cognitive prior combines ByteDance’s Seed-1.8, Retrieval-Augmented Generation, Chain-of-Thought reasoning, Qwen3.5-VL-generated instructions, and SAM3 grounding to infer functional parts and grasp types; and a temporal prior fuses SAM3 mask tracking with DINOv3 cross-frame features for rigid-motion updates. On DexGraspNet 2.0, spatial reasoning reaches an average 88.4% success rate across three hands and clutter levels. Functional inference reaches 94.0% functional-part accuracy and 87.0% grasp-type accuracy, while language-guided grasping achieves 81.3% on GraspClutter6D and 86.0% on GraspNet-1Billion. The policy is trained from 4.4M simulated grasp trials and transfers across two- to five-finger hands. In real-world tests using NVIDIA hardware, it achieves 83.3% success across 510 trials and 89.7% under human-induced dynamic disturbances, with online updates at 5 Hz. The key result is modularity: new perception or reasoning capabilities can be added without retraining the physical grasp policy.

Original abstract

This paper proposes AdaRoboVLG, a task-adaptive Vision-Language-Grasp (VLG) framework that supports generalizable grasp synthesis across different robotic hands. Unlike existing VLG methods that tightly couple foundation models with end-to-end grasp policies, AdaRoboVLG learns an efficient generalizable base policy that generates and evaluates physically feasible grasp candidates through explicit kinematic mapping and force-closure-based stability estimation, while offloading task-dependent understanding to specialized foundation-model modules. These modules provide composable priors that are integrated into the grasp synthesis process, enabling contextually adaptive grasp synthesis without retraining the underlying grasp policy. Through extensive simulation and real-world experiments, we demonstrate that (i) the base policy exhibits efficient learning and strong cross-hand generalization, (ii) the framework effectively incorporates spatial, cognitive, and temporal priors to address three representative grasping challenges without compromising grasp synthesis performance compared to state-of-the-art methods, and (iii) these priors can operate jointly to enable functional grasping in cluttered and dynamic environments. These results indicate that decoupling physical grasp synthesis from task-dependent understanding provides a scalable paradigm for robotic grasping, allowing future advances in foundation models to be directly translated into improved grasp capabilities without redesigning or retraining the underlying grasp policy. Supplementary videos are available at https://adarobovlg.github.io/

Read the original paper

More in Robotics

Browse all 50 papers →
02Robotics

Rolling-WAM: World Action Models with Rolling Imagination

Yinghua Zhou, Junjie Ye, Yiqi Zhao, Hao Dong, Celina Shiyu Wang, Ruohai Ge, Tingyi Yang, Basile Van Hoorick, Gaurav Sukhatme, Vitor Guizilini, Yue Wang

Rolling-WAM keeps future robot actions partially imagined and refined over time, making world-model-based manipulation replan 4.5 times faster.

Read analysis
03Robotics

Training-free Behavior Cloning

Maximilian Adang, Timothy Chen, Lars Osterberg, Aiden Swann, Mac Schwager

A fast, training-free robot controller reuses and corrects demonstration trajectories to deliver traceable behavior at real-time speeds.

Read analysis