NTH

TacBPM: A Tactile-conditioned Behavior Prior Model for Dexterous Reorientation

AuthorsJie Yin, Wanli Xing, Zeyuan Zhao, Xuezhou Zhu, Zhijie Deng, Kaifeng Zhang

AffiliationsSharpa Robotics

September 21, 2026 2 min read
Watch on YouTube
The one-line take

TacBPM gives dexterous robots a tactile-aware library of reusable motion skills so they can reorient unfamiliar objects more reliably and with less trial-and-error.

Key results

8
Sphere specialists

Number of multi-scale sphere teachers distilled into the tactile behavior prior

16-D
Latent action dimension

Gaussian latent action space used for residual downstream control

60.38%
Unseen-object average success

TacBPM success across unseen sphere scales and additional object shapes

70.00%
Multi-object success

TacBPM success on the mixed complex-object reorientation benchmark

97.36%
Staples-marker success

Arm-hand Grasp-to-AnyPose success for the unseen staples marker

What the paper found

TacBPM is a tactile-conditioned behavior prior for dexterous object reorientation, addressing a weakness of direct joint-space reinforcement learning used in earlier systems, including OpenAI’s dexterous manipulation work: contact-rich finger gaits are difficult to rediscover when object geometry or sensing changes. The method trains 8 sphere specialists with PPO in Isaac Sim and Isaac Lab, covering scale factors from 0.3 to 1.0, then distills them into a variational latent controller. Its 192-dimensional tactile-proprioceptive history conditions a 16-D Gaussian latent prior, while a decoder converts residual latent commands into a 22-DoF hand action. During downstream learning, the prior remains frozen and PPO explores around its contact-stable latent mean, with decoder finetuning adapting behavior to new geometries. On unseen sphere scales and non-spherical objects, TacBPM reaches 60.38% average success, compared with 40.05% for the no-tactile variant; on the mixed Multi reorientation benchmark, it achieves 70.00% success versus 1.00% for raw-action PPO. The same paradigm extends to arm-hand Grasp-to-AnyPose: on the unseen staples marker, success rises from 0.49% with raw-action PPO to 97.36% with TacBPM. Six-axis simulation, 20-hertz hardware control, and qualitative sim-to-real demonstrations show that tactile conditioning helps select contact-aware rolling, regrasping, and stabilization strategies, although command-switch transients and substantial geometry shifts remain failure modes.

Original abstract

Dexterous in-hand manipulation requires policies that coordinate high-DoF hand joints through intermittent, contact-rich interaction. Beyond target-orientation tracking, such policies must discover finger gaits that preserve object stability while adapting to geometry, anisotropy, pose, contact, and sensing changes. We propose \method, a tactile-conditioned behavior prior model for dexterous reorientation. \method distills multi-scale sphere specialists into a latent controller and lets downstream policies reuse the fixed tactile prior through residual latent actions, reducing renewed exploration from raw joint commands. The prior conditions on tactile-proprioceptive history so latent behavior reflects the current hand-object interaction. We evaluate arbitrary-pose transfer across anisotropic objects, commanded-axis rotation, and an arm-hand Grasp-to-AnyPose task in which the robot must grasp, lift, transport, and reach goal poses for novel tool geometries and generalized placements. Extensive experiments demonstrate that the proposed method accelerates training and enables stable policies where matched raw-action PPO remains near failure, with successful sim-to-real transfer in in-hand and arm-hand tasks.

Read the original paper

More in Robotics

Browse all 50 papers →
02Robotics

Rolling-WAM: World Action Models with Rolling Imagination

Yinghua Zhou, Junjie Ye, Yiqi Zhao, Hao Dong, Celina Shiyu Wang, Ruohai Ge, Tingyi Yang, Basile Van Hoorick, Gaurav Sukhatme, Vitor Guizilini, Yue Wang

Rolling-WAM keeps future robot actions partially imagined and refined over time, making world-model-based manipulation replan 4.5 times faster.

Read analysis
03Robotics

Training-free Behavior Cloning

Maximilian Adang, Timothy Chen, Lars Osterberg, Aiden Swann, Mac Schwager

A fast, training-free robot controller reuses and corrects demonstration trajectories to deliver traceable behavior at real-time speeds.

Read analysis