NTH

$\mathcal{N}_0$-Foundation: Towards the Age of Tactile Intelligence

AuthorsNeoteAI Team, Fudan TEAI Team

September 5, 2026 2 min read
Watch on YouTube
The one-line take

A large-scale tactile foundation platform that combines sensors, 30,000 hours of visuo-tactile data, transferable representations, and benchmarks for robot manipulation.

Key results

30000
NeoData interaction time

Total hours of synchronized visual-tactile demonstrations in NeoData.

5000
OpenNeoData subset

Hours in the released open-source subset.

0.066
NeoForce force MAE

Held-out force-field reconstruction MAE after adding latent prediction.

26.5%
π0.5 NeoReal success

Average binary success rate across the 10 NeoReal tasks.

47.5%
NeoForce progressive score

Average NeoReal progressive score when NeoForce conditions the π0.5 action expert.

What the paper found

N0-Foundation proposes a tactile-centered foundation for embodied manipulation, combining camera-based tactile hardware, the N0-TacUMI handheld interface, synchronized data collection, the NeoForce representation, and the NeoReal and NeoSim benchmarks. Its core dataset, NeoData, contains more than 30,000 hours of visual-tactile demonstrations across six embodiments and 450 tasks, including 1.4M episodes, 8B RGB frames, and 10B tactile frames; the released OpenNeoData subset contributes 5,000 hours for open research. NeoForce converts sensor-specific tactile images into dense three-axis force fields representing shear and pressure, then uses a shared transformer with masked reconstruction, temporal latent prediction, and cross-modal alignment. Adding latent prediction reduces force-field reconstruction MAE from 0.070 to 0.066 and RMSE from 0.095 to 0.089, while contact localization remains nearly unchanged. On NeoReal, the vision-language-action model π0.5 achieves a 26.5% average success rate, rising from a 38.1% to a 47.5% average progressive score when NeoForce conditions its action expert. On NeoSim, π0.5 reaches a 45.8% mean success rate, ahead of LingBot-VA at 32.1%. The experiments show that physical force representations are more transferable than device-specific tactile appearance, although contact-rich bimanual manipulation remains difficult for current policies, including Xiaomi-Robotics-0. The dataset annotation pipeline also uses Gemini-3.5-Flash for hierarchical task labeling, with signal-based temporal boundaries.

Original abstract

We present $\mathcal{N}_0$-Foundation, a paradigm for tactile-enabled embodied manipulation, which integrates tactile sensing hardware, large-scale multimodal data, tactile representation learning, and standardized evaluation. First, we engineer the infrastructure for scalable data collection, including a vision-based tactile sensor, a tactile Universal Manipulation Interface (UMI), and a synchronized visuo-tactile data collection system supporting both robot embodiments and UMI-based demonstrations. Leveraging this infrastructure, we construct NeoData, which contains more than 30000 hours of synchronized visual and tactile demonstrations, spanning six embodiments, 450 tasks, and billions of paired RGB and tactile frames collected through a mixture of real-robot teleoperation and UMI-based demonstrations. To facilitate open research, we further release OpenNeoData, a 5000-hour open-source subset of NeoData. The dataset addresses a central limitation of existing manipulation corpora, critical for deformable-object manipulation, precise assembly, delicate force control, and sustained surface interaction. Capitalizing on the large-scale, heterogeneous tactile measurements, we propose NeoForce, a visuo-tactile representation model that learn transferable tactile representations across different sensor designs. To enable systematic evaluation of tactile embodied models built upon our infrastructure, datasets and tactile representations, we further propose a comprehensive benchmark, which combines the real-world NeoReal suite and the simulated NeoSim suite for standardized evaluation. Experiments across both suites show that policies benefit from the physical contact state rather than from the device-specific appearance of the tactile signal. We release the dataset, the representation, and the benchmark, aiming at supporting future work on tactile-enabled embodied manipulation.

Read the original paper

More in Embodied AI

Browse all 48 papers →
01Embodied Ai

GroundingPI: A Grounding Foundation Model towards Physical Intelligence with Visual Primitives

Qize Yu, Lianrui Fan, Boyu Chen, Jiaqi Liang, Xini Ding, Yue Chen, Zetian Song, Yuran Wang, Yi Zou, Kaixuan Wang, Tianxing Chen, Wenxuan Song, Bohan Zhou, Mingleyang Li, Siqiao Huang, Yuqi Ye, Caigao Jiang, Wei Wei, Ruihai Wu, Hang Zhang, Yixiao Ge, Shuchang Zhou, Shilong Liu, Xianming Liu, Ping Luo, Shiyu Huang

GroundingPI argues that fast, precise visual grounding should be the perceptual foundation for capable robots and autonomous vehicles.

Read analysis