Deform360: A Massive Multi-view Visuotactile Dataset for Deformable World Models
AuthorsHongyu Li, Wanjia Fu, Xiaoyan Cong, Zekun Li, Binghao Huang, Hanxiao Jiang, Xintong He, Yiqing Liang, Rao Fu, Tao Lu, Srinath Sridhar, Kevin A. Smith, George Konidaris, Yunzhu Li
Resources
Deform360 is a big real-world dataset that helps robots learn how squishy objects move by combining many camera views, tactile sensing, and dense motion tracking.
Key results
daily-life deformable objects in Deform360
total manipulation episodes recorded
synchronized surround-view cameras
total multi-view video frames
visual-to-tactile contact prediction accuracy across 36 views
warped point cloud error with visuotactile tracking
What the paper found
Deform360 is a large-scale visuotactile benchmark for deformable world models, built by Brown University, Columbia University, and MIT to study whether 2D video generators or explicit 3D particle models better capture real manipulation physics. The dataset contains 198 daily-life objects, 1,980 interaction sequences, 41 synchronized surround-view cameras, and tactile grippers, totaling 23.3 million frames and 215.7 hours of recordings. Its markerless annotation pipeline combines per-frame 3D Gaussian Splatting, multi-view 2D point tracking with CoTracker3, 3D back-projection, RANSAC fusion, and tactile regularization to recover dense particle trajectories under heavy occlusion. On reconstruction, the pipeline reaches 27.66 dB PSNR overall, and tactile refinement reduces warped-point-cloud Chamfer error from 1.41×10^-4 m2 to 2.71×10^-5 m2, a fivefold improvement. For contact prediction, a transformer model achieves 88.67% accuracy and 0.8909 F1 across 36 synchronized views. In world-model benchmarking, PhysTwin is strongest in low-data per-episode settings, while Cosmos Predict 2.5 2B, an NVIDIA foundation video model, generalizes better in zero-shot multi-object prediction, revealing a concrete trade-off between structural priors and scaling.
Original abstract
Predicting object dynamics (i.e., world modeling) is a fundamental challenge for robotic manipulation, and modeling deformable objects presents a particularly difficult case due to their high-dimensional state spaces and complex material properties. While current world models approach this through two distinct paradigms: learning the dynamics over the 2D pixel space or more explicit 3D geometric space. A systematic understanding of their relative strengths and limitations remains elusive due to the lack of diverse, large-scale real-world data. To address this, we present Deform360, a large-scale visuotactile dataset featuring 198 daily-life objects, 1,980 interaction sequences, and over 215 hours of observations from 41 surround-view cameras and bimanual tactile grippers to capture both global motion and contact-induced local deformations. Leveraging a novel markerless visuotactile 3D tracking pipeline to extract dense geometry and motion, we systematically evaluate current state-of-the-art world models, comparing 2D video models against 3D particle models. Finally, we provide a preliminary demonstration indicating the real-world applicability of our dataset by performing robot planning tasks on deformable objects. Our analysis reveals key insights into the trade-offs between structural priors and scalability, providing a solid benchmark for future research in generalizable deformable object-centric world modeling. Project website: https://deform360.lhy.xyz
Read the original paperMore in Embodied AI
Browse all 48 papers →GroundingPI: A Grounding Foundation Model towards Physical Intelligence with Visual Primitives
Qize Yu, Lianrui Fan, Boyu Chen, Jiaqi Liang, Xini Ding, Yue Chen, Zetian Song, Yuran Wang, Yi Zou, Kaixuan Wang, Tianxing Chen, Wenxuan Song, Bohan Zhou, Mingleyang Li, Siqiao Huang, Yuqi Ye, Caigao Jiang, Wei Wei, Ruihai Wu, Hang Zhang, Yixiao Ge, Shuchang Zhou, Shilong Liu, Xianming Liu, Ping Luo, Shiyu Huang
GroundingPI argues that fast, precise visual grounding should be the perceptual foundation for capable robots and autonomous vehicles.
MM-ABC: Towards Generalist Mobile Manipulation via Seeing, Coordinating and Imagining
Qiwei Liang, Guangyu Chen, Shaolong Zhu, Zikuan Xiao, Jinxuan Lu, Yifan Xie, Renjing Xu, Wenbo Ding, Tianxing Chen
MM-ABC is a generalist robot foundation model that helps mobile manipulators see their surroundings, coordinate arm and base motion, and imagine future actions for better performance.
Morphometric Imitation: From Morphology and Contact Aware Hand Retargeting to Sim-to-Real Visuomotor Policy
Tara Sadjadpour, Siming He, C. K. Wolfe, Haozhi Qi, Lea Wilken, S. Shankar Sastry, Claire Tomlin, Jitendra Malik
A three-stage system converts human hand demonstrations into robust, zero-shot real-robot dexterous manipulation policies across different hand morphologies.