FA-RDP: A Frequency-Adaptive Reactive Diffusion Policy for Contact-Rich Manipulation
AuthorsLifeng Zhuo, Wendi Chen, Han Xue, Shirun Tang, Jun Lv, Cewu Lu, Chuan Wen
Resources
FA-RDP lets robot policies explore diverse approaches before contact and react quickly to force feedback once manipulation becomes constrained.
Key results
Average success across Dual Box Flipping, Dual Switch Toggling, and Dual Button Pressing.
Average-success improvement in percentage points over the 51.7% ImplicitRDP baseline.
Low-frequency multi-step diffusion mode used for ambiguous pre-contact behavior.
One-step distilled mode used for reactive post-contact force control.
Inference latency remained below 30 ms for high-frequency execution.
What the paper found
FA-RDP, from researchers at Shanghai Jiao Tong University, the Shanghai Innovation Institute, and Noematrix, addresses a core problem in contact-rich manipulation: robots need diverse approach trajectories before contact but rapid force-feedback reactions afterward. Its shared visual-force Transformer uses frequency-adaptive positional encoding to generate both low-frequency and high-frequency action chunks, while a learned multimodality indicator selects multi-step diffusion at 10 Hz during ambiguous pre-contact motion and switches to a one-step distilled sampler at 30 Hz when contact constraints reduce ambiguity. The proposed Manifold Consistency Distillation, or MCD, trains the high-frequency model to predict actions directly on the robot action manifold while retaining DDPM residual supervision, avoiding unstable noise-, score-, or velocity-target distillation. On three real-world tasks—Dual Box Flipping, Dual Switch Toggling, and Dual Button Pressing—with 60 demonstrations and 20 trials per method, FA-RDP achieved an average success rate of 81.7%, improving on the strongest main baseline, ImplicitRDP at 51.7%, by 30.0 percentage points. Indicator-guided switching also outperformed always using the distilled high-frequency policy, which achieved 61.7%. Both modes span the same 1.6 s horizon, and high-frequency inference latency remained below 30 ms, demonstrating a practical balance between pre-contact multimodality and post-contact reactivity.
Original abstract
In contact-rich manipulation, action multimodality and reactivity dominate different stages of a single episode. Before contact, multiple trajectories might be equally valid, making it important to preserve diverse action modes. After contact, geometric constraints and force limits narrow the solution space, while successful execution demands rapid responses to force feedback. However, standard diffusion policies use a fixed inference frequency and sampling steps throughout the episode, forcing a fundamental compromise: low-frequency, multi-step sampling better preserves pre-contact multimodality but responds slowly to force feedback, whereas high-frequency sampling improves reactivity but tends to collapse distinct pre-contact modes. To resolve this tradeoff, we present FA-RDP, a frequency-adaptive reactive diffusion policy. A shared multi-frequency visual-force Transformer predicts action chunks at both low and high frequencies, while a learned multimodality indicator dynamically selects multi-step low-frequency sampling before contact and one-step high-frequency sampling as action ambiguity decreases. We further introduce Manifold Consistency Distillation (MCD), which reparameterizes the diffusion network to predict actions on the robot action manifold while retaining DDPM-based residual supervision. Experiments on three contact-rich manipulation tasks show that FA-RDP achieves the highest success rate while preserving diverse pre-contact trajectory modes. Code and videos are available at https://fa-rdp.github.io.
Read the original paperMore in Robotics
Browse all 50 papers →JAMB: Joint Action-Motion Diffusion for Bimanual Manipulation
Chuyang Xiao, Peilin Meng, David Held
JAMB helps two robot arms coordinate by jointly imagining their future movements and the changing 3D scene before acting.
Rolling-WAM: World Action Models with Rolling Imagination
Yinghua Zhou, Junjie Ye, Yiqi Zhao, Hao Dong, Celina Shiyu Wang, Ruohai Ge, Tingyi Yang, Basile Van Hoorick, Gaurav Sukhatme, Vitor Guizilini, Yue Wang
Rolling-WAM keeps future robot actions partially imagined and refined over time, making world-model-based manipulation replan 4.5 times faster.
Training-free Behavior Cloning
Maximilian Adang, Timothy Chen, Lars Osterberg, Aiden Swann, Mac Schwager
A fast, training-free robot controller reuses and corrects demonstration trajectories to deliver traceable behavior at real-time speeds.