NTH

From Raw Experience to Skill Consumption: A Systematic Study of Model-Generated Agent Skills

AuthorsZisu Huang, Jingwen Xu, Yifan Yang, Ziyang Gong, Qihao Yang, Muzhao Tian, Xiaohua Wang, Changze Lv, Xuemei Gao, Qi Dai, Bei Liu, Kai Qiu, Xue Yang, Dongdong Chen, Xiaoqing Zheng, Chong Luo

June 11, 2026 2 min read
Watch on YouTube
The one-line take

This paper studies how AI agents turn raw experience into reusable skills, finds that these skills can help but also hurt in surprising ways, and proposes a meta-skill to make skill extraction more reliable.

Key results

75%
Positive-transfer rate

Share of extractor–target pairs with ∆ > 0

25%
Negative-transfer rate

Share of extractor–target pairs with ∆ < 0

46.4%
Judge accuracy

Unguided GPT-5.4 pairwise skill preference accuracy

73.8%
Judge accuracy with validated rubric

Pairwise skill preference accuracy after rubric guidance

1.55
Meta-skill gain

Average downstream improvement in percentage points from the validated rubric

-0.59
Plausibility-rubric effect

Average downstream change in percentage points from the naive plausibility rubric

What the paper found

This Microsoft Research, Fudan University, and Shanghai Jiao Tong University study systematically evaluates model-generated agent skills across the full lifecycle of raw experience generation, skill extraction, and skill consumption, using five domains, six target models, and five extractor models. Across ALFWorld, SpreadsheetBench, SWE-bench-Verified, SEAL-0, and BFCL-v4, extracted skills improve downstream performance in 75% of extractor–target pairs, but 25% still show negative transfer, with ALFWorld the most fragile domain at 47% negative cases. The paper shows that extraction quality is not a simple function of model strength: on SpreadsheetBench, Gemini-3.1-Flash-Lite achieves the highest extraction efficacy even though GPT-5.4 has the strongest baseline task score, and the same skill can transfer very differently across targets. A deeper analysis finds that the success-to-failure mix in the experience pool is domain-dependent, that skill format itself is irrelevant, and that an LLM judge’s unconstrained pairwise preference is only 46.4% accurate and falls to 15.8% on the hardest pairs. By contrast, a validated three-dimension rubric—Failure Mechanism Encoding, Actionable Specificity, and High-Risk Action Blacklist—raises judge accuracy to 73.8% and, when converted into a meta-skill inserted into the extractor prompt, improves average downstream performance by 1.55 percentage points while the naive plausibility rubric hurts by 0.59 points.

Original abstract

Language agents increasingly improve by reusing \emph{skills} -- structured procedural artifacts distilled from past experience. In particular, \emph{domain-level} and \emph{model-generated} skills are especially promising. They offer fast adaptation within a domain by encoding domain-specific recurring procedures, and they scale beyond labor-intensive hand-crafting. However, while extraction methods continue to proliferate, understanding remains limited, with no comprehensive study spanning the full skill lifecycle -- \textbf{experience generation}, \textbf{skill extraction}, and \textbf{skill consumption} -- to ask whether such skills actually work, when they work, and what makes them succeed or fail. To close this gap, we build a utility-grounded evaluation framework that provides systematic experimental results across extractors and target agents, covering five diverse agentic task domains. We find that model-generated skills are beneficial on average but exhibit non-trivial negative transfer, and that neither extractors nor targets behave uniformly. A model can be a strong extractor yet a weak consumer, or vice versa, with skill utility independent of model scale or baseline task strength. To explain these patterns, we then dissect each lifecycle stage in depth, analyzing how experience composition shapes skill quality, what properties characterize useful skills, and how the same skill transfers across different consumers. Finally, we translate these findings into a concrete \emph{meta-skill} that guides skill extraction toward the features tied to actual utility, which consistently improves skill quality across domains and substantially reduces negative transfer.

Read the original paper

More in AI Agents

Browse all 56 papers →
01Agent

LEGO-Anything: Coding Agents for 3D Scene Reconstruction

Xirui Li, Peng Shi, Mingwen Dong, Sheng Zhang, Zhuoyan Xu, Dongkyu Lee, Shuaichen Chang, Yi Xiang, Lin Pan, Jiarong Jiang

LEGO-Anything turns images into editable Blender programs through iterative coding agents, offering a promising but still imperfect route to reconstructable 3D worlds.

Read analysis
02Agent

MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution

Prithwish Jana, Mononito Goswami, Hao Liu, Xinyu Li, Langlin Huang, Zhehui Huang, Zhishen Huang, Patrick Blöbaum, Anoop Deoras, Purak Jain, Nikos Kanakaris, Sahika Genc

MILO uses teams of evolving AI agents to automatically discover better harnesses for long-horizon problem-solving systems.

Read analysis
03Agent

Self-Organizing Agent Teams Learn to Reason Together

Aneesh Pappu, Mirac Suzgun, Yongchan Kwon, Federico Bianchi, Batu El, Mykel J. Kochenderfer, Hancheng Cao, James Zou

This work trains AI agents to discover how to divide labor, challenge ideas, and combine reasoning so that teams can solve problems no individual agent could solve alone.

Read analysis