NTH

$N_0$-VTLA: Scaling Vision-Tactile-Language-Action Model with Latent Tactile Tokens

AuthorsNeoteAI Team, Fudan TEAI Team

August 4, 2026 2 min read
Watch on YouTube
The one-line take

N0-VTLA combines large-scale tactile learning with offline reinforcement learning to give robots substantially better control of contact-rich and deformable objects.

Key results

92.3%
Future-tactile retrieval

Top-1 accuracy for retrieving the matching future-tactile target from predicted latent tokens.

47.2%
NeoReal mean success

Average success across all nine real-robot NeoReal tasks.

29.4%
π0.5 NeoReal baseline

Average success of the strongest reported baseline on NeoReal.

63.8%
20-task simulation mean

Mean N0-VTLA success across the 20-task UniVTAC and NeoSim simulation suite.

95%
ALTER Towel Folding

Success rate of N0-VTLA trained with ALTER on a long-horizon real-robot task.

What the paper found

NeoteAI and the Fudan TEAI Team introduce N0-VTLA, a vision–tactile–language–action foundation model built on Physical Intelligence’s π0.5 and PaliGemma. Instead of inserting sparse tactile readings into the vision–language context, N0-VTLA encodes DINOv2 features from contact-difference images, predicts latent tactile tokens representing contact changes over a 50-step action chunk, and feeds those tokens directly to a flow-matching action expert. A three-stage training recipe grounds the tactile predictor, aligns it with action generation, and then fine-tunes the full policy. On held-out representation tests, predicted latents retrieve the matching future-tactile target with 92.3% top-1 accuracy, versus 3.2% chance. The model wins all nine NeoReal real-robot tasks, averaging 47.2% success compared with 29.4% for π0.5, and reaches 63.8% mean success across 20 UniVTAC and NeoSim simulation tasks, ahead of 44.0% for the strongest overall baseline. The authors also introduce ALTER, an advantage-conditioned offline reinforcement-learning method that combines duration-weighted stage progress, tactile-detected object-drop events, and human corrections to label stored deployment data. With N0-VTLA plus ALTER, success reaches 95% on Towel Folding, 80% on Bag Packing, and 75% on Cardboard Box Folding, demonstrating that predictive touch improves both contact-aware control and offline policy refinement.

Original abstract

We present $N_0$-VTLA, a vision-tactile-language-action (VTLA) foundation model capable of (1) fine-grained contact-rich manipulation with tactile perception and tactile-feedback control, and (2) offline policy improvement from stored deployment data. Building on current vision-based backbones, we propose a training recipe for tactile integration consisting of visuo-tactile pre-training, staged tactile-pathway integration, and advantage-conditioned offline policy improvement. During pre-training, the policy learns broad contact priors from NeoData, our large-scale visuo-tactile robot dataset; to our knowledge, $N_0$-VTLA is the first VTLA model pretrained on tactile data at scale. During post-training, we augment the policy with a predictive tactile pathway that distills the contact patterns learned at scale into the fine motion adjustments required by downstream tactile-centric manipulation. For offline policy improvement, we introduce ALTER, an advantage-conditioned offline reinforcement learning method that converts relative progress and trajectory-event comparisons into binary advantage labels for policy training on a fixed deployment corpus, further improving task-specific learning on contact-rich skills such as deformable object manipulation. Across contact-rich benchmarks, $N_0$-VTLA outperforms strong baselines by wide margins: it wins all nine real-robot NeoReal tasks and reaches 63.8% mean success on a twenty-task simulation suite, against 44.0% for the strongest baseline. $N_0$-VTLA policies trained with ALTER reach 75-95% success on three long-horizon real-robot tasks. These results lay a foundation for versatile tactile-driven manipulation policies.

Read the original paper

More in Embodied AI

Browse all 48 papers →
01Embodied Ai

GroundingPI: A Grounding Foundation Model towards Physical Intelligence with Visual Primitives

Qize Yu, Lianrui Fan, Boyu Chen, Jiaqi Liang, Xini Ding, Yue Chen, Zetian Song, Yuran Wang, Yi Zou, Kaixuan Wang, Tianxing Chen, Wenxuan Song, Bohan Zhou, Mingleyang Li, Siqiao Huang, Yuqi Ye, Caigao Jiang, Wei Wei, Ruihai Wu, Hang Zhang, Yixiao Ge, Shuchang Zhou, Shilong Liu, Xianming Liu, Ping Luo, Shiyu Huang

GroundingPI argues that fast, precise visual grounding should be the perceptual foundation for capable robots and autonomous vehicles.

Read analysis