$N_0$-VTLA: Scaling Vision-Tactile-Language-Action Model with Latent Tactile Tokens
AuthorsNeoteAI Team, Fudan TEAI Team
Resources
N0-VTLA combines large-scale tactile learning with offline reinforcement learning to give robots substantially better control of contact-rich and deformable objects.
Key results
Top-1 accuracy for retrieving the matching future-tactile target from predicted latent tokens.
Average success across all nine real-robot NeoReal tasks.
Average success of the strongest reported baseline on NeoReal.
Mean N0-VTLA success across the 20-task UniVTAC and NeoSim simulation suite.
Success rate of N0-VTLA trained with ALTER on a long-horizon real-robot task.
What the paper found
NeoteAI and the Fudan TEAI Team introduce N0-VTLA, a vision–tactile–language–action foundation model built on Physical Intelligence’s π0.5 and PaliGemma. Instead of inserting sparse tactile readings into the vision–language context, N0-VTLA encodes DINOv2 features from contact-difference images, predicts latent tactile tokens representing contact changes over a 50-step action chunk, and feeds those tokens directly to a flow-matching action expert. A three-stage training recipe grounds the tactile predictor, aligns it with action generation, and then fine-tunes the full policy. On held-out representation tests, predicted latents retrieve the matching future-tactile target with 92.3% top-1 accuracy, versus 3.2% chance. The model wins all nine NeoReal real-robot tasks, averaging 47.2% success compared with 29.4% for π0.5, and reaches 63.8% mean success across 20 UniVTAC and NeoSim simulation tasks, ahead of 44.0% for the strongest overall baseline. The authors also introduce ALTER, an advantage-conditioned offline reinforcement-learning method that combines duration-weighted stage progress, tactile-detected object-drop events, and human corrections to label stored deployment data. With N0-VTLA plus ALTER, success reaches 95% on Towel Folding, 80% on Bag Packing, and 75% on Cardboard Box Folding, demonstrating that predictive touch improves both contact-aware control and offline policy refinement.
Original abstract
We present $N_0$-VTLA, a vision-tactile-language-action (VTLA) foundation model capable of (1) fine-grained contact-rich manipulation with tactile perception and tactile-feedback control, and (2) offline policy improvement from stored deployment data. Building on current vision-based backbones, we propose a training recipe for tactile integration consisting of visuo-tactile pre-training, staged tactile-pathway integration, and advantage-conditioned offline policy improvement. During pre-training, the policy learns broad contact priors from NeoData, our large-scale visuo-tactile robot dataset; to our knowledge, $N_0$-VTLA is the first VTLA model pretrained on tactile data at scale. During post-training, we augment the policy with a predictive tactile pathway that distills the contact patterns learned at scale into the fine motion adjustments required by downstream tactile-centric manipulation. For offline policy improvement, we introduce ALTER, an advantage-conditioned offline reinforcement learning method that converts relative progress and trajectory-event comparisons into binary advantage labels for policy training on a fixed deployment corpus, further improving task-specific learning on contact-rich skills such as deformable object manipulation. Across contact-rich benchmarks, $N_0$-VTLA outperforms strong baselines by wide margins: it wins all nine real-robot NeoReal tasks and reaches 63.8% mean success on a twenty-task simulation suite, against 44.0% for the strongest baseline. $N_0$-VTLA policies trained with ALTER reach 75-95% success on three long-horizon real-robot tasks. These results lay a foundation for versatile tactile-driven manipulation policies.
Read the original paperMore in Embodied AI
Browse all 48 papers →GroundingPI: A Grounding Foundation Model towards Physical Intelligence with Visual Primitives
Qize Yu, Lianrui Fan, Boyu Chen, Jiaqi Liang, Xini Ding, Yue Chen, Zetian Song, Yuran Wang, Yi Zou, Kaixuan Wang, Tianxing Chen, Wenxuan Song, Bohan Zhou, Mingleyang Li, Siqiao Huang, Yuqi Ye, Caigao Jiang, Wei Wei, Ruihai Wu, Hang Zhang, Yixiao Ge, Shuchang Zhou, Shilong Liu, Xianming Liu, Ping Luo, Shiyu Huang
GroundingPI argues that fast, precise visual grounding should be the perceptual foundation for capable robots and autonomous vehicles.
MM-ABC: Towards Generalist Mobile Manipulation via Seeing, Coordinating and Imagining
Qiwei Liang, Guangyu Chen, Shaolong Zhu, Zikuan Xiao, Jinxuan Lu, Yifan Xie, Renjing Xu, Wenbo Ding, Tianxing Chen
MM-ABC is a generalist robot foundation model that helps mobile manipulators see their surroundings, coordinate arm and base motion, and imagine future actions for better performance.
Morphometric Imitation: From Morphology and Contact Aware Hand Retargeting to Sim-to-Real Visuomotor Policy
Tara Sadjadpour, Siming He, C. K. Wolfe, Haozhi Qi, Lea Wilken, S. Shankar Sastry, Claire Tomlin, Jitendra Malik
A three-stage system converts human hand demonstrations into robust, zero-shot real-robot dexterous manipulation policies across different hand morphologies.