Grasp, Handover, Rotate: Bimanual Object Reorientation via Compositional Diffusion and Energy-Based Optimization
AuthorsWun Lam Yeung, Wenjun Liu, Yui Cheung Yu, Zhengyan Lambo Qin, Qijin She, Heng Li, Ziqi Wang, Ping Tan
Resources
A diffusion-and-energy-based robot planner helps two arms grasp, hand over, rotate, and place objects more reliably and smoothly.
Key results
Number of pick-and-place reorientation tasks across easy, medium, and hard scenes.
Collision-free trajectory success rate on the benchmark.
Reduction reported for BiCompoDiff-Full relative to the NoEBM configuration.
Absolute success-rate gain: 81.7% versus 58.3%.
Reduction in joint displacement for the T-shaped brick experiment.
Reduction in joint displacement for the cup experiment.
What the paper found
Researchers at The Hong Kong University of Science and Technology and Shenzhen Loop Area Institute introduce BiCompoDiff, a unified framework for bimanual object reorientation: one UR12e arm grasps and hands over an object, while the other regrasps and places it in a new pose. BiCompoDiff combines the pretrained GraspGen 6-DoF diffusion model with differentiable energy-based guidance for collision avoidance, handover feasibility, regrasp safety, and joint-space smoothness. During reverse diffusion, annealed MCMC refines grasp, handover, and placement poses, while the learned SubnetIK model supplies fast gradients for inverse-kinematics feasibility; cuRobo then generates minimum-jerk trajectories. On a benchmark of 60 simulated tasks spanning easy, medium, and hard clutter, the full system reduced joint displacement by 37% and achieved a 63.3% success rate versus 41.7% without planning-energy guidance. Against an adapted ReorientBot baseline, it reached 81.7% success versus 58.3%, a 23.4% absolute improvement, while also reducing joint displacement by 12.5%. Real-robot tests transferred the method to two UR12e arms with Robotiq grippers: joint displacement improved by 59.7% for a T-shaped brick and 46.4% for a cup. The results demonstrate that compositional diffusion can optimize grasp quality and downstream motion jointly rather than relying on costly sample-and-filter pipelines. The project involves HKUST and Shenzhen Loop Area Institute and acknowledges support from the BYD-HKUST Joint Lab.
Original abstract
Bimanual object reorientation - picking an object, handing it over between two arms, and placing it in a desired target pose - is valuable when direct placement from the initial grasp is infeasible due to collisions, kinematic constraints, or poor final orientation. However, achieving this under multiple competing objectives remains challenging. We introduce BiCompoDiff, a compositional diffusion and energy-based framework that jointly optimizes grasp selection, handover, regrasp, and motion planning under multiple constraints. By combining a pretrained grasp diffusion model with bimanual planning energy-based models (EBMs), our method injects gradient guidance during reverse diffusion to enforce collision avoidance, trajectory smoothness (via differentiable inverse kinematics), handover feasibility, and regrasp safety. Annealed MCMC sampling further refines grasp poses over the composite energy landscape. Experiments across diverse simulated household reorientation tasks demonstrate that BiCompoDiff achieves over 20% higher success rates and up to 37% smoother trajectories (measured by joint displacement) compared to strong sampling-based baselines. Real-world validation confirms effective sim-to-real transfer and robust performance on challenging scenes.
Read the original paperMore in Robotics
Browse all 50 papers →JAMB: Joint Action-Motion Diffusion for Bimanual Manipulation
Chuyang Xiao, Peilin Meng, David Held
JAMB helps two robot arms coordinate by jointly imagining their future movements and the changing 3D scene before acting.
Rolling-WAM: World Action Models with Rolling Imagination
Yinghua Zhou, Junjie Ye, Yiqi Zhao, Hao Dong, Celina Shiyu Wang, Ruohai Ge, Tingyi Yang, Basile Van Hoorick, Gaurav Sukhatme, Vitor Guizilini, Yue Wang
Rolling-WAM keeps future robot actions partially imagined and refined over time, making world-model-based manipulation replan 4.5 times faster.
Training-free Behavior Cloning
Maximilian Adang, Timothy Chen, Lars Osterberg, Aiden Swann, Mac Schwager
A fast, training-free robot controller reuses and corrects demonstration trajectories to deliver traceable behavior at real-time speeds.