NTH

Playful Agentic Robot Learning

AuthorsJunyi Zhang, Jiaxin Ge, Hanjun Yoo, Letian Fu, Zihan Yang, Yaowei Liu, Raj Saravanan, Shaofeng Yin, Justin Yu, Dantong Niu, Zirui Wang, Roei Herzig, Ken Goldberg, Yutong Bai, David M. Chan, Ion Stoica, Angjoo Kanazawa, Jiahui Lei, Haiwen Feng, Trevor Darrell

June 19, 2026 2 min read
Watch on YouTube
The one-line take

This paper teaches robots to play first, so they can build reusable code-based skills that help them solve future tasks better.

Key results

20.6
LIBERO-PRO gain

Average success improvement over CaP-Agent0

17.0
MolmoSpaces gain

Average success improvement over CaP-Agent0

8.9
RoboSuite transfer gain

Cross-environment success improvement when LIBERO-PRO skills are plugged into CaP-Agent0

8.8
Real-world gain

Real-world success improvement when LIBERO-PRO skills are plugged into CaP-Agent0

3.3%
MolmoSpaces real-world average

Baseline real-world success before adding RAT S skills

25.0%
MolmoSpaces real-world average with skills

Real-world success after adding RAT S skills

What the paper found

Playful Agentic Robot Learning introduces RAT S, a multi-agent Code-as-Policy framework in which a robot does not wait for downstream instructions but first engages in self-directed play to accumulate reusable skills. During play, a Task Proposer uses a Goldilocks novelty-and-learnability criterion to select exploratory manipulation goals, and an Execution Team writes, verifies, retries, and diagnoses robot programs before successful behaviors are distilled into a persistent skill library. The method is evaluated on LIBERO-PRO and MolmoSpaces, where play-learned skills raise held-out success by 20.6 percentage points and 17.0 percentage points over CaP-Agent0, respectively, while a random-play baseline yields only marginal gains. The frozen library also transfers across environments: plugging LIBERO-PRO skills into CaP-Agent0 improves RoboSuite by 8.9 points and real-world manipulation by 8.8 points, and MolmoSpaces-learned skills improve a separate real-world set from 3.3% to 25.0%. An ablation shows that curiosity-driven play matters more than random exploration, and that play-time skill acquisition composes best with the full RAT S verification loop rather than extra test-time retries. The main technical contribution is a practical mechanism for converting autonomous interaction traces into named, reusable code skills for future manipulation tasks.

Original abstract

Current agentic robot systems can write executable Code-as-Policy programs, observe feedback, and revise behavior across multiple attempts, but they remain largely task-driven: reusable skills are acquired only after explicit instructions. We study Playful Agentic Robot Learning, where an embodied coding agent uses self-directed play as a continual skill-learning stage before downstream tasks arrive. We introduce RATs, Robotics Agent Teams designed for play-time skill acquisition. During play, RATs proposes novel yet learnable exploratory tasks, plans and executes robot-code policies, verifies intermediate progress, diagnoses failures, retries with dense, step-level feedback, and distills successful executions into a persistent code skill library. At test time, the agent reuses relevant skills from this frozen library to help solve new tasks. Experiments in LIBERO-PRO and MolmoSpaces show that play-learned skills improve held-out downstream tasks over no-play and random-play baselines, with 20.6 and 17.0 percentage-point gains over CaP-Agent0 on LIBERO-PRO and MolmoSpaces, respectively. Moreover, the learned skills can be plugged into other inference-time Code-as-Policy agents by simply retrieving them into the context, improving RoboSuite and real-world transfer by 8.9 and 8.8 points, respectively, without finetuning the underlying model.

Read the original paper

More in Robotics

Browse all 50 papers →
02Robotics

Rolling-WAM: World Action Models with Rolling Imagination

Yinghua Zhou, Junjie Ye, Yiqi Zhao, Hao Dong, Celina Shiyu Wang, Ruohai Ge, Tingyi Yang, Basile Van Hoorick, Gaurav Sukhatme, Vitor Guizilini, Yue Wang

Rolling-WAM keeps future robot actions partially imagined and refined over time, making world-model-based manipulation replan 4.5 times faster.

Read analysis
03Robotics

Training-free Behavior Cloning

Maximilian Adang, Timothy Chen, Lars Osterberg, Aiden Swann, Mac Schwager

A fast, training-free robot controller reuses and corrects demonstration trajectories to deliver traceable behavior at real-time speeds.

Read analysis