InSight: Self-Guided Skill Acquisition via Steerable VLAs
AuthorsMaggie Wang, Lars Osterberg, Stephen Tian, Ola Shorinwa, Jiajun Wu, Mac Schwager
Resources
InSight lets robot policies teach themselves new manipulation skills by breaking demonstrations into reusable primitives and then autonomously practicing, labeling, and adding missing skills to their training data.
Key results
Success rate after 246 acquired primitive rollouts in simulation
Total rollouts needed to reach the reported success rate
Trials to obtain 20 successful twist primitives on xArm
Trials to obtain 20 successful pour primitives on xArm
End-to-end success on the 14-primitive composition task
Evaluation trials succeeded for the acquired sweeping primitive
What the paper found
InSight, from Stanford University and Princeton University, reframes vision-language-action learning by making a policy steerable at the primitive level and then using a vision-language model, Gemini 3 Flash, to discover and acquire missing primitives without new human demonstrations. The system first auto-segments teleoperated data by aligning VLM-generated primitive plans with gripper transitions, end-effector motion, and visual cues, then fine-tunes a π0.5 VLA with LoRA and a learned progress channel. In the second stage, the VLM decomposes a novel task, flags primitive gaps, proposes single-axis parameters for each gap, and filters successful rollouts with oracle checks before retraining the VLA on the newly acquired primitives. On simulation block flipping in LIBERO, 150 pick-and-place demonstrations were segmented into over 700 primitive episodes, and InSight reached 75% success after 246 acquired primitive rollouts, while an SAC baseline never completed a flip under the same rollout budget. On hardware with a 6DoF UFactory xArm, InSight acquired 20 successful twist primitives in 23 trials and 20 successful pour primitives in 31 trials, then achieved 92% success on twisting, 96% on pouring, and 80% on a 14-primitive twist-then-pour composition, compared with 32%, 16%, and 4% for CaP-X and 0% on the π0.5 baseline. The unified policy retained 100% success on original pick-and-place skills, and InSight also learned sweeping from scooping demonstrations, succeeding in 5/5 trials.
Original abstract
Vision-language-action (VLA) models can learn manipulation skills from demonstrations, but their capabilities are bounded by the skills in the training data. We present InSight, a framework that unlocks autonomous skill acquisition by rendering VLAs steerable at the primitive-action level (e.g., "move gripper to the bowl", "lift upward", "pour the bottle"). InSight consists of two primary stages: (1) an automated segmentation pipeline that partitions demonstrations into labeled primitives via VLM plan decomposition and end-effector poses to enable VLA primitive steerability, and (2) a VLM-guided data flywheel that identifies missing primitives required to accomplish a novel task, autonomously attempts demonstrations of the missing primitives with VLM-proposed low-level control, and automatically labels, stores, and integrates successful demonstrations into the VLA training set. We evaluate InSight across simulation and real-world manipulation tasks, including block flipping, drawer closing, sweeping, twisting, and pouring, without any human demonstrations of these target skills. Once learned, these primitives can be composed to execute novel, long-horizon tasks without additional human demonstrations. Our findings demonstrate that primitive steerability provides a practical foundation for continual skill acquisition in VLA policies. Project website: https://insight-vla.github.io.
Read the original paperMore in Embodied AI
Browse all 48 papers →GroundingPI: A Grounding Foundation Model towards Physical Intelligence with Visual Primitives
Qize Yu, Lianrui Fan, Boyu Chen, Jiaqi Liang, Xini Ding, Yue Chen, Zetian Song, Yuran Wang, Yi Zou, Kaixuan Wang, Tianxing Chen, Wenxuan Song, Bohan Zhou, Mingleyang Li, Siqiao Huang, Yuqi Ye, Caigao Jiang, Wei Wei, Ruihai Wu, Hang Zhang, Yixiao Ge, Shuchang Zhou, Shilong Liu, Xianming Liu, Ping Luo, Shiyu Huang
GroundingPI argues that fast, precise visual grounding should be the perceptual foundation for capable robots and autonomous vehicles.
MM-ABC: Towards Generalist Mobile Manipulation via Seeing, Coordinating and Imagining
Qiwei Liang, Guangyu Chen, Shaolong Zhu, Zikuan Xiao, Jinxuan Lu, Yifan Xie, Renjing Xu, Wenbo Ding, Tianxing Chen
MM-ABC is a generalist robot foundation model that helps mobile manipulators see their surroundings, coordinate arm and base motion, and imagine future actions for better performance.
Morphometric Imitation: From Morphology and Contact Aware Hand Retargeting to Sim-to-Real Visuomotor Policy
Tara Sadjadpour, Siming He, C. K. Wolfe, Haozhi Qi, Lea Wilken, S. Shankar Sastry, Claire Tomlin, Jitendra Malik
A three-stage system converts human hand demonstrations into robust, zero-shot real-robot dexterous manipulation policies across different hand morphologies.