NTH

Deform360: A Massive Multi-view Visuotactile Dataset for Deformable World Models

AuthorsHongyu Li, Wanjia Fu, Xiaoyan Cong, Zekun Li, Binghao Huang, Hanxiao Jiang, Xintong He, Yiqing Liang, Rao Fu, Tao Lu, Srinath Sridhar, Kevin A. Smith, George Konidaris, Yunzhu Li

July 13, 2026 2 min read
Watch on YouTube
The one-line take

Deform360 is a big real-world dataset that helps robots learn how squishy objects move by combining many camera views, tactile sensing, and dense motion tracking.

Key results

198
objects

daily-life deformable objects in Deform360

1980
interaction sequences

total manipulation episodes recorded

41
camera views

synchronized surround-view cameras

23.3M
frames

total multi-view video frames

88.67%
contact accuracy

visual-to-tactile contact prediction accuracy across 36 views

2.71e-5
tactile Chamfer error

warped point cloud error with visuotactile tracking

What the paper found

Deform360 is a large-scale visuotactile benchmark for deformable world models, built by Brown University, Columbia University, and MIT to study whether 2D video generators or explicit 3D particle models better capture real manipulation physics. The dataset contains 198 daily-life objects, 1,980 interaction sequences, 41 synchronized surround-view cameras, and tactile grippers, totaling 23.3 million frames and 215.7 hours of recordings. Its markerless annotation pipeline combines per-frame 3D Gaussian Splatting, multi-view 2D point tracking with CoTracker3, 3D back-projection, RANSAC fusion, and tactile regularization to recover dense particle trajectories under heavy occlusion. On reconstruction, the pipeline reaches 27.66 dB PSNR overall, and tactile refinement reduces warped-point-cloud Chamfer error from 1.41×10^-4 m2 to 2.71×10^-5 m2, a fivefold improvement. For contact prediction, a transformer model achieves 88.67% accuracy and 0.8909 F1 across 36 synchronized views. In world-model benchmarking, PhysTwin is strongest in low-data per-episode settings, while Cosmos Predict 2.5 2B, an NVIDIA foundation video model, generalizes better in zero-shot multi-object prediction, revealing a concrete trade-off between structural priors and scaling.

Original abstract

Predicting object dynamics (i.e., world modeling) is a fundamental challenge for robotic manipulation, and modeling deformable objects presents a particularly difficult case due to their high-dimensional state spaces and complex material properties. While current world models approach this through two distinct paradigms: learning the dynamics over the 2D pixel space or more explicit 3D geometric space. A systematic understanding of their relative strengths and limitations remains elusive due to the lack of diverse, large-scale real-world data. To address this, we present Deform360, a large-scale visuotactile dataset featuring 198 daily-life objects, 1,980 interaction sequences, and over 215 hours of observations from 41 surround-view cameras and bimanual tactile grippers to capture both global motion and contact-induced local deformations. Leveraging a novel markerless visuotactile 3D tracking pipeline to extract dense geometry and motion, we systematically evaluate current state-of-the-art world models, comparing 2D video models against 3D particle models. Finally, we provide a preliminary demonstration indicating the real-world applicability of our dataset by performing robot planning tasks on deformable objects. Our analysis reveals key insights into the trade-offs between structural priors and scalability, providing a solid benchmark for future research in generalizable deformable object-centric world modeling. Project website: https://deform360.lhy.xyz

Read the original paper

More in Embodied AI

Browse all 48 papers →
01Embodied Ai

GroundingPI: A Grounding Foundation Model towards Physical Intelligence with Visual Primitives

Qize Yu, Lianrui Fan, Boyu Chen, Jiaqi Liang, Xini Ding, Yue Chen, Zetian Song, Yuran Wang, Yi Zou, Kaixuan Wang, Tianxing Chen, Wenxuan Song, Bohan Zhou, Mingleyang Li, Siqiao Huang, Yuqi Ye, Caigao Jiang, Wei Wei, Ruihai Wu, Hang Zhang, Yixiao Ge, Shuchang Zhou, Shilong Liu, Xianming Liu, Ping Luo, Shiyu Huang

GroundingPI argues that fast, precise visual grounding should be the perceptual foundation for capable robots and autonomous vehicles.

Read analysis