Rapid On-Robot Learning for Dynamic Manipulation Skills: Robot Juggling
AuthorsTaeyoon Lee, Chunpeng Wang, Christopher G. Atkeson, Alfred A. Rizzi, Nicolas Rojas
Resources
A bimanual robot learns five juggling patterns safely in under five minutes by refining imperfect prior models with its own real-world experience.
Key results
The robot learned cascade, tennis, half-shower, shower, and box three-ball patterns.
Safety ablation evaluated 7,578 real-world task-level planner queries.
Without the Mutually Reachable Set, 89.0% of planner solutions were unsafe.
EfficientTAM achieved 0.01 second inference time on an NVIDIA RTX A6000.
What the paper found
This paper presents a bimanual robot that learns dynamic manipulation directly on hardware by combining an imperfect prior with regularized memory-based learning. For each throw, a k-nearest-neighbor memory retrieves relevant physical experience, fits a locally weighted linear model regularized toward the prior, and updates the landing command without broad unsafe exploration. A precomputed Mutually Reachable Set constrains joint position, velocity, acceleration, and jerk transitions, while Ruckig generates feasible real-time trajectories. Using onboard vision based on EfficientTAM, FastGICP, and synchronized RGB-D sensing, the AthenaZero platform learned and composed five three-ball patterns—cascade, tennis, half-shower, shower, and box—under 5 minutes of real-world interaction, despite a prior model that could not complete a single cycle. Across 7,578 planner queries, removing the safety constraint produced 89.0% unsafe solutions, whereas the constrained planner kept 100% of solutions safe by construction. EfficientTAM ran at 0.01 second inference time on an NVIDIA RTX A6000, supporting low-latency perception without task-specific retraining. The cascade pattern was learned reliably after seven resets, and experience transferred across shared skills when the robot transitioned among patterns. Shower and box remained less deterministic because their short flight times prevented visual feedback from correcting catches, highlighting the need for tactile or proprioceptive feedback.
Original abstract
We present an online learning framework that enables a bimanual robot to acquire diverse juggling patterns directly on physical hardware within minutes, even with a significant sim2real gap. One of the most important lessons from this work is that a model, even when far from reality, can be extremely useful for learning. This motivates a central philosophy of our approach: learning should build upon the robot's current knowledge rather than replace it. Our regularized memory-based learning puts this principle into practice by learning a local model from accumulated experience while retaining the global prior model to extrapolate where experience is sparse. This enables efficient and stable online learning from each new experience without resorting to uninformed exploration over a vast space of possible behaviors. Equally important to continual on-robot learning is safety, allowing the robot to repeatedly practice and improve in the real world. We construct a mutually reachable set that allows safe transitions between successive throws and catches, without driving either arm into a state from which its next action would require violating the robot's joint or actuator limits. Together, these ideas enable a bimanual robot with multi-fingered hands and onboard vision to safely learn and compose five canonical three-ball juggling patterns, including cascade, tennis, half-shower, shower, and box, within less than 5 minutes of real-world interaction. More broadly, this work points toward robots that build upon imperfect prior knowledge and continually refine their behavior through their own real-world experience.
Read the original paperMore in Robotics
Browse all 50 papers →JAMB: Joint Action-Motion Diffusion for Bimanual Manipulation
Chuyang Xiao, Peilin Meng, David Held
JAMB helps two robot arms coordinate by jointly imagining their future movements and the changing 3D scene before acting.
Rolling-WAM: World Action Models with Rolling Imagination
Yinghua Zhou, Junjie Ye, Yiqi Zhao, Hao Dong, Celina Shiyu Wang, Ruohai Ge, Tingyi Yang, Basile Van Hoorick, Gaurav Sukhatme, Vitor Guizilini, Yue Wang
Rolling-WAM keeps future robot actions partially imagined and refined over time, making world-model-based manipulation replan 4.5 times faster.
Training-free Behavior Cloning
Maximilian Adang, Timothy Chen, Lars Osterberg, Aiden Swann, Mac Schwager
A fast, training-free robot controller reuses and corrects demonstration trajectories to deliver traceable behavior at real-time speeds.