SkillOpt: Executive Strategy for Self-Evolving Agent Skills
AuthorsYifan Yang, Ziyang Gong, Weiquan Huang, Qihao Yang, Ziwei Zhou, Zisu Huang, Yan Li, Xuemei Gao, Qi Dai, Bei Liu, Kai Qiu, Yuqing Yang, Dongdong Chen, Xue Yang, Chong Luo
Resources
SkillOpt treats an agent’s skill as something you can train like a model parameter, using controlled text edits and validation to make AI agents steadily improve without extra runtime cost.
Key results
SkillOpt is best or tied-best on all benchmark-model-harness combinations
Average improvement over no skill in direct chat
Average improvement over no skill inside Codex
Average improvement over no skill inside Claude Code
Codex-trained skill transferred to Claude Code
What the paper found
SkillOpt, from Microsoft with collaborators at Shanghai Jiao Tong University, Tongji University, and Fudan University, reframes agent-skill improvement as a controlled text-space optimization problem: a frozen target model executes tasks while a separate optimizer model edits one reusable skill document through bounded add/delete/replace operations, held-out validation gating, rejected-edit buffering, and an epoch-wise slow/meta update. Across six benchmarks, seven target models, and three execution harnesses—direct chat, OpenAI Codex, and Anthropic Claude Code—it is best or tied-best in all 52 evaluated cells, outperforming human-written, one-shot LLM, Trace2Skill, TextGrad, GEPA, and EvoSkill skills. On GPT–5.5, the average gain over no skill is +23.5 points in direct chat, +24.8 in Codex, and +19.1 in Claude Code, with large procedural gains such as 41.8 to 80.7 on SpreadsheetBench, 33.1 to 72.1 on OfficeQA, and 37.6 to 66.9 on LiveMathematicianBench. The learned artifacts stay compact, only 379 to 1,995 tokens, and require just 1 to 4 accepted edits. Transfer is also a central result: a Codex-trained SpreadsheetBench skill transfers to Claude Code with a +59.7-point gain, and a cross-benchmark OlympiadBench skill improves Omni-MATH by +3.7 on GPT–5.4. Ablations show that removing the rejected-edit buffer drops SpreadsheetBench from 77.5 to 72.9, while removing both meta skill and slow update collapses it to 55.0, confirming that SkillOpt’s stability comes from validation-gated bounded updates rather than unconstrained rewriting.
Original abstract
Agent skills today are hand-crafted, generated one-shot, or evolved through loosely controlled self-revision, none of which behaves like a deep-learning optimizer for the skill, and none of which reliably improves over its starting point under feedback. We argue the skill should instead be trained as the external state of a frozen agent, with the same discipline that makes weight-space optimization reproducible. SkillOpt is, to our knowledge, the first systematic controllable text-space optimizer for agent skills: a separate optimizer model turns scored rollouts into bounded add/delete/replace edits on a single skill document, and an edit is accepted only when it strictly improves a held-out validation score. A textual learning-rate budget, rejected-edit buffer, and epoch-wise slow/meta update make skill training stable while adding zero inference-time model calls at deployment. Across six benchmarks, seven target models, and three execution harnesses (direct chat, Codex, Claude Code), SkillOpt is best or tied on all 52 evaluated (model, benchmark, harness) cells and beats every per-cell competitor among human, one-shot LLM, Trace2Skill, TextGrad, GEPA, and EvoSkill skills. On GPT-5.5 it lifts the average no-skill accuracy by +23.5 points in direct chat, by +24.8 inside the Codex agentic loop, and by +19.1 inside Claude Code. Transfer experiments further show that optimized skill artifacts retain value when moved across model scales, between Codex and Claude Code execution environments, and to a nearby math benchmark without further optimization.
Read the original paperMore in AI Agents
Browse all 56 papers →LEGO-Anything: Coding Agents for 3D Scene Reconstruction
Xirui Li, Peng Shi, Mingwen Dong, Sheng Zhang, Zhuoyan Xu, Dongkyu Lee, Shuaichen Chang, Yi Xiang, Lin Pan, Jiarong Jiang
LEGO-Anything turns images into editable Blender programs through iterative coding agents, offering a promising but still imperfect route to reconstructable 3D worlds.
MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution
Prithwish Jana, Mononito Goswami, Hao Liu, Xinyu Li, Langlin Huang, Zhehui Huang, Zhishen Huang, Patrick Blöbaum, Anoop Deoras, Purak Jain, Nikos Kanakaris, Sahika Genc
MILO uses teams of evolving AI agents to automatically discover better harnesses for long-horizon problem-solving systems.
Self-Organizing Agent Teams Learn to Reason Together
Aneesh Pappu, Mirac Suzgun, Yongchan Kwon, Federico Bianchi, Batu El, Mykel J. Kochenderfer, Hancheng Cao, James Zou
This work trains AI agents to discover how to divide labor, challenge ideas, and combine reasoning so that teams can solve problems no individual agent could solve alone.