NTH

Sample-Efficient Learning from Agent Experience

AuthorsChenhui Gou, Haoqin Tu, Yunhao Fang, Jianfei Cai, Hamid Rezatofighi

July 27, 2026 2 min read
Watch on YouTube
The one-line take

Experience Distillation lets agents retain much of what they learn from trial and error without repeatedly paying the cost of interacting with the environment.

Key results

749
Curated SWE tasks

Software-engineering benchmark size

51.4%
SWE Experience Distillation pass@1

Average pass@1 after distillation

64.8%
ICL gain retained on SWE

Fraction of in-context-learning improvement retained

3.8%
SFT gain retained on SWE

Gain retained by direct supervised fine-tuning

9.6
SWE sample-efficiency improvement

Times fewer environment samples than PPO

57.2
TaleSuite sample-efficiency improvement

Times fewer environment samples than GRPO

What the paper found

Researchers from Monash University and ByteDance Seed introduce Experience Distillation, a method for consolidating an agent’s trial-and-error learning into model weights without collecting additional environment data. The approach uses one-step branched rollouts: at histories already present in recorded trajectories, an experience-conditioned teacher generates only its next action, and a student learns that action through sampled-token next-token prediction. Experience preprocessing, enhanced teacher reasoning, and branch packing make the procedure practical without a learned world model, avoiding compounding synthetic-rollout errors. Across 749 curated software-engineering tasks and six TaleSuite text-adventure games, Experience Distillation reaches 51.4% pass@1 on software engineering and retains 64.8% of the improvement that in-context learning provides, while direct supervised fine-tuning retains only 3.8%. On TaleSuite, it achieves an average normalized score of 43.8 and matches classical reinforcement-learning performance with 9.6× fewer environment samples on software engineering and 57.2× fewer on the games. The method also generalizes to 494 out-of-distribution software tasks and supports continual collect-and-distill cycles, suggesting a workflow in which agents first adapt rapidly in context and then permanently internalize useful experience.

Original abstract

Real-world agent learning is often constrained by costly environment interactions, such as running time-consuming experiments or obtaining human feedback. In-context learning offers a highly sample-efficient way for agents to learn from their own interaction histories, but its gains disappear once that experience is removed from the context. Separately, context distillation provides a mechanism for internalizing contextual information into model weights. However, applying it to agents' interaction histories without sacrificing environment sample efficiency remains underexplored. We term this problem Experience Distillation and develop an implementation that requires no further environment interaction beyond the collected experience. Experiments on 749 curated software-engineering tasks and six text-adventure games show that it retains at least 64.8\% of the gains from in-context learning across both domains, whereas direct supervised fine-tuning on the collected experience recovers only 3.8\%. Compared with classical reinforcement-learning baselines, in-context learning from trial-and-error experience followed by Experience Distillation matches their performance with at least \(9.6\times\) fewer environment samples.

Read the original paper

More in AI Agents

Browse all 56 papers →
01Agent

LEGO-Anything: Coding Agents for 3D Scene Reconstruction

Xirui Li, Peng Shi, Mingwen Dong, Sheng Zhang, Zhuoyan Xu, Dongkyu Lee, Shuaichen Chang, Yi Xiang, Lin Pan, Jiarong Jiang

LEGO-Anything turns images into editable Blender programs through iterative coding agents, offering a promising but still imperfect route to reconstructable 3D worlds.

Read analysis
02Agent

MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution

Prithwish Jana, Mononito Goswami, Hao Liu, Xinyu Li, Langlin Huang, Zhehui Huang, Zhishen Huang, Patrick Blöbaum, Anoop Deoras, Purak Jain, Nikos Kanakaris, Sahika Genc

MILO uses teams of evolving AI agents to automatically discover better harnesses for long-horizon problem-solving systems.

Read analysis
03Agent

Self-Organizing Agent Teams Learn to Reason Together

Aneesh Pappu, Mirac Suzgun, Yongchan Kwon, Federico Bianchi, Batu El, Mykel J. Kochenderfer, Hancheng Cao, James Zou

This work trains AI agents to discover how to divide labor, challenge ideas, and combine reasoning so that teams can solve problems no individual agent could solve alone.

Read analysis