NTH

JAMB: Joint Action-Motion Diffusion for Bimanual Manipulation

AuthorsChuyang Xiao, Peilin Meng, David Held

AffiliationsRobotics Institute, Carnegie Mellon University, Pittsburgh, PA 15213, USA · University of Michigan, Ann Arbor, MI 48109, USA

October 3, 2026 2 min read
Watch on YouTube
The one-line take

JAMB helps two robot arms coordinate by jointly imagining their future movements and the changing 3D scene before acting.

Key results

83.4%
RoboTwin 2.0 overall success

Average success across 16 simulation tasks.

23.9
Simulation improvement

Percentage-point improvement over the strongest baseline.

85.6%
Real-world success

Average success across three bimanual robot tasks.

50.0
Real-world improvement over DP3

Percentage-point improvement over the action-only DP3 baseline.

7.6
3D track ADE

Average displacement error in millimeters for moving patches.

17.9%
Easy-to-Hard generalization

Success after training on Easy scenes and testing on Hard scenes without fine-tuning.

What the paper found

JAMB is a diffusion policy for coordinated bimanual manipulation that jointly denoises robot action chunks and future 3D point tracks, rather than predicting actions alone or using future geometry as fixed auxiliary conditioning. A shared Diffusion Transformer combines DINOv2 visual patches, proprioception, actions, and patch-aligned 3D trajectories, while 4D RoPE encodes world-space position and physical time so each arm’s proposed motion can refine the predicted scene evolution throughout denoising. On 16 RoboTwin 2.0 simulation tasks, JAMB reaches an 83.4% average success rate, exceeding the strongest baseline by 23.9 percentage points. On three real-world tasks using two xArm manipulators, a ZED Mini camera, and a Meta VR headset for demonstrations, it achieves 85.6%, compared with 35.6% for DP3 and 64.4% for GAP—improvements of 50.0 and 21.2 percentage points. The full model predicts moving-point trajectories with 7.6 millimeters ADE and 10.0 millimeters FDE, and joint denoising improves ablation success to 81.7%, versus 76.3% for track-conditioned action diffusion. For Easy-to-Hard transfer without fine-tuning, JAMB reaches 17.9%, compared with 4.3% for GAP, indicating stronger robustness to clutter and unseen object appearances. Real-world track supervision uses CoTracker3 and FoundationStereo, while only the generated actions are executed; the 3D tracks function as an internal geometric prediction that guides coordination.

Original abstract

Coordinated bimanual manipulation is challenging because the motion of either arm can alter the shared 3D scene and thereby affect the other arm. Yet most diffusion policies generate actions without explicitly modeling these future geometric consequences, while predictive variants typically use future state only as auxiliary supervision or fixed conditioning. We address this limitation by proposing JAMB, a diffusion policy that jointly denoises bimanual actions and future 3D point tracks. By allowing action and track hypotheses to evolve together within a shared Transformer, each can inform and refine the other throughout denoising. We further ground multimodal representations in a shared spatiotemporal coordinate system to facilitate geometry-aware interaction during joint denoising. We evaluate JAMB on diverse bimanual manipulation tasks in RoboTwin 2.0 and on a real-world robot, comparing it with action-only policies and alternative future-prediction approaches spanning different state representations and learning objectives. Across 16 simulation tasks, JAMB achieves an average success rate of 83.4%, outperforming the strongest baseline by 23.9 percentage points. On three real-world tasks, it outperforms the action-only and auxiliary geometry prediction methods by 50.0 and 21.2 percentage points, respectively. Beyond these performance gains, JAMB shows stronger generalization to cluttered scenes and out-of-distribution backgrounds than the evaluated baselines. Together, these results demonstrate the effectiveness of our joint action-motion modeling framework for coordinated bimanual manipulation. Our project website is available at https://jam-bimanual.github.io/

Read the original paper

More in Robotics

Browse all 50 papers →
01Robotics

Rolling-WAM: World Action Models with Rolling Imagination

Yinghua Zhou, Junjie Ye, Yiqi Zhao, Hao Dong, Celina Shiyu Wang, Ruohai Ge, Tingyi Yang, Basile Van Hoorick, Gaurav Sukhatme, Vitor Guizilini, Yue Wang

Rolling-WAM keeps future robot actions partially imagined and refined over time, making world-model-based manipulation replan 4.5 times faster.

Read analysis
02Robotics

Training-free Behavior Cloning

Maximilian Adang, Timothy Chen, Lars Osterberg, Aiden Swann, Mac Schwager

A fast, training-free robot controller reuses and corrects demonstration trajectories to deliver traceable behavior at real-time speeds.

Read analysis