TacBPM: A Tactile-conditioned Behavior Prior Model for Dexterous Reorientation
AuthorsJie Yin, Wanli Xing, Zeyuan Zhao, Xuezhou Zhu, Zhijie Deng, Kaifeng Zhang
AffiliationsSharpa Robotics
Resources
TacBPM gives dexterous robots a tactile-aware library of reusable motion skills so they can reorient unfamiliar objects more reliably and with less trial-and-error.
Key results
Number of multi-scale sphere teachers distilled into the tactile behavior prior
Gaussian latent action space used for residual downstream control
TacBPM success across unseen sphere scales and additional object shapes
TacBPM success on the mixed complex-object reorientation benchmark
Arm-hand Grasp-to-AnyPose success for the unseen staples marker
What the paper found
TacBPM is a tactile-conditioned behavior prior for dexterous object reorientation, addressing a weakness of direct joint-space reinforcement learning used in earlier systems, including OpenAI’s dexterous manipulation work: contact-rich finger gaits are difficult to rediscover when object geometry or sensing changes. The method trains 8 sphere specialists with PPO in Isaac Sim and Isaac Lab, covering scale factors from 0.3 to 1.0, then distills them into a variational latent controller. Its 192-dimensional tactile-proprioceptive history conditions a 16-D Gaussian latent prior, while a decoder converts residual latent commands into a 22-DoF hand action. During downstream learning, the prior remains frozen and PPO explores around its contact-stable latent mean, with decoder finetuning adapting behavior to new geometries. On unseen sphere scales and non-spherical objects, TacBPM reaches 60.38% average success, compared with 40.05% for the no-tactile variant; on the mixed Multi reorientation benchmark, it achieves 70.00% success versus 1.00% for raw-action PPO. The same paradigm extends to arm-hand Grasp-to-AnyPose: on the unseen staples marker, success rises from 0.49% with raw-action PPO to 97.36% with TacBPM. Six-axis simulation, 20-hertz hardware control, and qualitative sim-to-real demonstrations show that tactile conditioning helps select contact-aware rolling, regrasping, and stabilization strategies, although command-switch transients and substantial geometry shifts remain failure modes.
Original abstract
Dexterous in-hand manipulation requires policies that coordinate high-DoF hand joints through intermittent, contact-rich interaction. Beyond target-orientation tracking, such policies must discover finger gaits that preserve object stability while adapting to geometry, anisotropy, pose, contact, and sensing changes. We propose \method, a tactile-conditioned behavior prior model for dexterous reorientation. \method distills multi-scale sphere specialists into a latent controller and lets downstream policies reuse the fixed tactile prior through residual latent actions, reducing renewed exploration from raw joint commands. The prior conditions on tactile-proprioceptive history so latent behavior reflects the current hand-object interaction. We evaluate arbitrary-pose transfer across anisotropic objects, commanded-axis rotation, and an arm-hand Grasp-to-AnyPose task in which the robot must grasp, lift, transport, and reach goal poses for novel tool geometries and generalized placements. Extensive experiments demonstrate that the proposed method accelerates training and enables stable policies where matched raw-action PPO remains near failure, with successful sim-to-real transfer in in-hand and arm-hand tasks.
Read the original paperMore in Robotics
Browse all 50 papers →JAMB: Joint Action-Motion Diffusion for Bimanual Manipulation
Chuyang Xiao, Peilin Meng, David Held
JAMB helps two robot arms coordinate by jointly imagining their future movements and the changing 3D scene before acting.
Rolling-WAM: World Action Models with Rolling Imagination
Yinghua Zhou, Junjie Ye, Yiqi Zhao, Hao Dong, Celina Shiyu Wang, Ruohai Ge, Tingyi Yang, Basile Van Hoorick, Gaurav Sukhatme, Vitor Guizilini, Yue Wang
Rolling-WAM keeps future robot actions partially imagined and refined over time, making world-model-based manipulation replan 4.5 times faster.
Training-free Behavior Cloning
Maximilian Adang, Timothy Chen, Lars Osterberg, Aiden Swann, Mac Schwager
A fast, training-free robot controller reuses and corrects demonstration trajectories to deliver traceable behavior at real-time speeds.