ChainSplat: A Physics-Inspired Screw-Theoretic Model for Learning Deformable Linear Object Dynamics from Multi-View RGB Videos
AuthorsSeungyeon Kim, Noémie Jaquier
Resources
ChainSplat combines physics-inspired articulated modeling with Gaussian splatting to reconstruct and control deformable objects such as cables from ordinary multi-view video.
Key results
Trajectories collected across three ropes and two interaction scenarios.
Each ChainSplat model uses eight screw joints and nine links.
Intersection over Union for the 20 cm rope in free-space evaluation.
RGB rendering PSNR for the 20 cm rope in free-space evaluation.
Average training time in minutes on an NVIDIA A100.
Total average inference time in seconds, including rollout and rendering.
What the paper found
ChainSplat is a physics-inspired digital-twin framework for deformable linear objects such as ropes and cables, learned solely from synchronized multi-view RGB videos rather than depth. It models each object as an open chain of rigid links connected by revolute joints, using screw-theoretic product-of-exponentials kinematics, recursive Newton–Euler dynamics, torsional damping, gravity, and tabletop friction; link geometry and appearance are represented with differentiable 3D Gaussian Splatting. A two-stage optimization first recovers link-aware Gaussians and joint trajectories by minimizing rendered-versus-observed RGB error, using Meta’s Segment Anything Model for object masks, then identifies masses, damping, and friction by matching simulated and recovered trajectories. Experiments used 24 trajectories across three ropes, with ChainSplat configured as 8 screw joints and 9 links, and compared it with PGND and PhysTwin. In free space for the 20 cm rope, it achieved an IoU of 0.704 and PSNR of 33.65, while training required 17.6 minutes and total inference, including rollout and rendering, required 2.41 seconds on an NVIDIA A100. On an RTX 4090, single-camera RGB state estimation ran at approximately 2–3 Hz, and model-based target hitting achieved average error below 1 cm after dynamics identification. The result is a compact, physically interpretable state representation that supports RGB-based state estimation, external-force inference, and gradient-based trajectory optimization for real-world DLO manipulation.
Original abstract
Identifying the underlying dynamics and 3D geometry of deformable linear objects (DLOs), such as cables, ropes, and hoses, is essential for accurate robotic manipulation, but remains challenging due to their high-dimensional configuration spaces and diverse behaviors arising from varying material properties. Existing methods often rely on multi-stage pipelines and auxiliary depth inputs, which are prone to errors under dynamic interactions, while their high-dimensional state representations make model-based control computationally expensive. In this paper, we introduce ChainSplat, a physics-inspired framework that jointly learns the 3D geometry, appearance, kinematics, and dynamics of DLOs solely from multi-view RGB videos. ChainSplat represents a DLO as an open-chain structure of rigid links connected by revolute joints, yielding an analytic, screw-theoretic model with a compact state representation parameterized by joint configurations. By integrating this formulation with Gaussian splatting, ChainSplat jointly recovers DLO dynamics, kinematics-aware 3D geometry, and appearance, while enabling high-fidelity RGB rendering from arbitrary states. Through real-world experiments, we demonstrate that ChainSplat achieves state-of-the-art performance in dynamics predictions, 3D geometry reconstruction, and RGB rendering across dynamic interactions. ChainSplat further enables real-time state and force estimation, as well as accurate model-based trajectory optimization, highlighting its practical utility for real-world robotic manipulation of DLOs. Accompanying source code and video are available at: https://chainsplat.github.io.
Read the original paperMore in Robotics
Browse all 50 papers →JAMB: Joint Action-Motion Diffusion for Bimanual Manipulation
Chuyang Xiao, Peilin Meng, David Held
JAMB helps two robot arms coordinate by jointly imagining their future movements and the changing 3D scene before acting.
Rolling-WAM: World Action Models with Rolling Imagination
Yinghua Zhou, Junjie Ye, Yiqi Zhao, Hao Dong, Celina Shiyu Wang, Ruohai Ge, Tingyi Yang, Basile Van Hoorick, Gaurav Sukhatme, Vitor Guizilini, Yue Wang
Rolling-WAM keeps future robot actions partially imagined and refined over time, making world-model-based manipulation replan 4.5 times faster.
Training-free Behavior Cloning
Maximilian Adang, Timothy Chen, Lars Osterberg, Aiden Swann, Mac Schwager
A fast, training-free robot controller reuses and corrects demonstration trajectories to deliver traceable behavior at real-time speeds.