Playful Agentic Robot Learning
AuthorsJunyi Zhang, Jiaxin Ge, Hanjun Yoo, Letian Fu, Zihan Yang, Yaowei Liu, Raj Saravanan, Shaofeng Yin, Justin Yu, Dantong Niu, Zirui Wang, Roei Herzig, Ken Goldberg, Yutong Bai, David M. Chan, Ion Stoica, Angjoo Kanazawa, Jiahui Lei, Haiwen Feng, Trevor Darrell
Resources
This paper teaches robots to play first, so they can build reusable code-based skills that help them solve future tasks better.
Key results
Average success improvement over CaP-Agent0
Average success improvement over CaP-Agent0
Cross-environment success improvement when LIBERO-PRO skills are plugged into CaP-Agent0
Real-world success improvement when LIBERO-PRO skills are plugged into CaP-Agent0
Baseline real-world success before adding RAT S skills
Real-world success after adding RAT S skills
What the paper found
Playful Agentic Robot Learning introduces RAT S, a multi-agent Code-as-Policy framework in which a robot does not wait for downstream instructions but first engages in self-directed play to accumulate reusable skills. During play, a Task Proposer uses a Goldilocks novelty-and-learnability criterion to select exploratory manipulation goals, and an Execution Team writes, verifies, retries, and diagnoses robot programs before successful behaviors are distilled into a persistent skill library. The method is evaluated on LIBERO-PRO and MolmoSpaces, where play-learned skills raise held-out success by 20.6 percentage points and 17.0 percentage points over CaP-Agent0, respectively, while a random-play baseline yields only marginal gains. The frozen library also transfers across environments: plugging LIBERO-PRO skills into CaP-Agent0 improves RoboSuite by 8.9 points and real-world manipulation by 8.8 points, and MolmoSpaces-learned skills improve a separate real-world set from 3.3% to 25.0%. An ablation shows that curiosity-driven play matters more than random exploration, and that play-time skill acquisition composes best with the full RAT S verification loop rather than extra test-time retries. The main technical contribution is a practical mechanism for converting autonomous interaction traces into named, reusable code skills for future manipulation tasks.
Original abstract
Current agentic robot systems can write executable Code-as-Policy programs, observe feedback, and revise behavior across multiple attempts, but they remain largely task-driven: reusable skills are acquired only after explicit instructions. We study Playful Agentic Robot Learning, where an embodied coding agent uses self-directed play as a continual skill-learning stage before downstream tasks arrive. We introduce RATs, Robotics Agent Teams designed for play-time skill acquisition. During play, RATs proposes novel yet learnable exploratory tasks, plans and executes robot-code policies, verifies intermediate progress, diagnoses failures, retries with dense, step-level feedback, and distills successful executions into a persistent code skill library. At test time, the agent reuses relevant skills from this frozen library to help solve new tasks. Experiments in LIBERO-PRO and MolmoSpaces show that play-learned skills improve held-out downstream tasks over no-play and random-play baselines, with 20.6 and 17.0 percentage-point gains over CaP-Agent0 on LIBERO-PRO and MolmoSpaces, respectively. Moreover, the learned skills can be plugged into other inference-time Code-as-Policy agents by simply retrieving them into the context, improving RoboSuite and real-world transfer by 8.9 and 8.8 points, respectively, without finetuning the underlying model.
Read the original paperMore in Robotics
Browse all 50 papers →JAMB: Joint Action-Motion Diffusion for Bimanual Manipulation
Chuyang Xiao, Peilin Meng, David Held
JAMB helps two robot arms coordinate by jointly imagining their future movements and the changing 3D scene before acting.
Rolling-WAM: World Action Models with Rolling Imagination
Yinghua Zhou, Junjie Ye, Yiqi Zhao, Hao Dong, Celina Shiyu Wang, Ruohai Ge, Tingyi Yang, Basile Van Hoorick, Gaurav Sukhatme, Vitor Guizilini, Yue Wang
Rolling-WAM keeps future robot actions partially imagined and refined over time, making world-model-based manipulation replan 4.5 times faster.
Training-free Behavior Cloning
Maximilian Adang, Timothy Chen, Lars Osterberg, Aiden Swann, Mac Schwager
A fast, training-free robot controller reuses and corrects demonstration trajectories to deliver traceable behavior at real-time speeds.