NTH

InSight: Self-Guided Skill Acquisition via Steerable VLAs

AuthorsMaggie Wang, Lars Osterberg, Stephen Tian, Ola Shorinwa, Jiajun Wu, Mac Schwager

June 26, 2026 2 min read
Watch on YouTube
The one-line take

InSight lets robot policies teach themselves new manipulation skills by breaking demonstrations into reusable primitives and then autonomously practicing, labeling, and adding missing skills to their training data.

Key results

75%
block flip success

Success rate after 246 acquired primitive rollouts in simulation

246
block flip rollout attempts

Total rollouts needed to reach the reported success rate

23
twist acquisition trials

Trials to obtain 20 successful twist primitives on xArm

31
pour acquisition trials

Trials to obtain 20 successful pour primitives on xArm

80%
twist-then-pour success

End-to-end success on the 14-primitive composition task

5/5
sweeping success

Evaluation trials succeeded for the acquired sweeping primitive

What the paper found

InSight, from Stanford University and Princeton University, reframes vision-language-action learning by making a policy steerable at the primitive level and then using a vision-language model, Gemini 3 Flash, to discover and acquire missing primitives without new human demonstrations. The system first auto-segments teleoperated data by aligning VLM-generated primitive plans with gripper transitions, end-effector motion, and visual cues, then fine-tunes a π0.5 VLA with LoRA and a learned progress channel. In the second stage, the VLM decomposes a novel task, flags primitive gaps, proposes single-axis parameters for each gap, and filters successful rollouts with oracle checks before retraining the VLA on the newly acquired primitives. On simulation block flipping in LIBERO, 150 pick-and-place demonstrations were segmented into over 700 primitive episodes, and InSight reached 75% success after 246 acquired primitive rollouts, while an SAC baseline never completed a flip under the same rollout budget. On hardware with a 6DoF UFactory xArm, InSight acquired 20 successful twist primitives in 23 trials and 20 successful pour primitives in 31 trials, then achieved 92% success on twisting, 96% on pouring, and 80% on a 14-primitive twist-then-pour composition, compared with 32%, 16%, and 4% for CaP-X and 0% on the π0.5 baseline. The unified policy retained 100% success on original pick-and-place skills, and InSight also learned sweeping from scooping demonstrations, succeeding in 5/5 trials.

Original abstract

Vision-language-action (VLA) models can learn manipulation skills from demonstrations, but their capabilities are bounded by the skills in the training data. We present InSight, a framework that unlocks autonomous skill acquisition by rendering VLAs steerable at the primitive-action level (e.g., "move gripper to the bowl", "lift upward", "pour the bottle"). InSight consists of two primary stages: (1) an automated segmentation pipeline that partitions demonstrations into labeled primitives via VLM plan decomposition and end-effector poses to enable VLA primitive steerability, and (2) a VLM-guided data flywheel that identifies missing primitives required to accomplish a novel task, autonomously attempts demonstrations of the missing primitives with VLM-proposed low-level control, and automatically labels, stores, and integrates successful demonstrations into the VLA training set. We evaluate InSight across simulation and real-world manipulation tasks, including block flipping, drawer closing, sweeping, twisting, and pouring, without any human demonstrations of these target skills. Once learned, these primitives can be composed to execute novel, long-horizon tasks without additional human demonstrations. Our findings demonstrate that primitive steerability provides a practical foundation for continual skill acquisition in VLA policies. Project website: https://insight-vla.github.io.

Read the original paper

More in Embodied AI

Browse all 48 papers →
01Embodied Ai

GroundingPI: A Grounding Foundation Model towards Physical Intelligence with Visual Primitives

Qize Yu, Lianrui Fan, Boyu Chen, Jiaqi Liang, Xini Ding, Yue Chen, Zetian Song, Yuran Wang, Yi Zou, Kaixuan Wang, Tianxing Chen, Wenxuan Song, Bohan Zhou, Mingleyang Li, Siqiao Huang, Yuqi Ye, Caigao Jiang, Wei Wei, Ruihai Wu, Hang Zhang, Yixiao Ge, Shuchang Zhou, Shilong Liu, Xianming Liu, Ping Luo, Shiyu Huang

GroundingPI argues that fast, precise visual grounding should be the perceptual foundation for capable robots and autonomous vehicles.

Read analysis