OpenSkill: Open-World Self-Evolution for LLM Agents
AuthorsZhiling Yan, Dingjie Song, Hanrong Zhang, Wei Liang, Yuxuan Zhang, Yutong Dai, Lifang He, Philip S. Yu, Ran Xu, Xiang Li, Lichao Sun
Resources
OpenSkill lets LLM agents teach themselves from open-web resources and self-generated practice tasks, even when no labeled target data or verifier exists.
Key results
OpenSkill overall score on SkillsBench with Opus 4.6
OpenSkill overall score on SkillsBench with GPT 5.2
Points above the strongest closed-world baseline on SkillsBench
Points above the strongest closed-world baseline on SkillsBench
Agreement with ground-truth evaluation outcomes on 84 cases
Coverage of ground-truth test intents by virtual tests
What the paper found
OpenSkill: Open-World Self-Evolution for LLM Agents, from Lehigh University, the University of Illinois Chicago, Salesforce AI Research, the University of British Columbia, Vector Institute, and Harvard Medical School, reframes agent improvement as a no-supervision problem: starting from only a task prompt, a base model, and open-world resources, the agent must invent both reusable skills and its own verification signal without seeing ground-truth tests. The framework bootstraps this loop in three stages: open-world retrieval from documentation, repositories, and the web; leakage-free skill evolution against self-built virtual tasks anchored to independently verifiable facts; and zero-shot deployment of the frozen skill artifact to the target agent. On SkillsBench, SocialMaze, and ScienceWorld, OpenSkill is the best automated method across all reported target-agent settings, including 43.6% overall pass rate on Opus 4.6 and 42.1% on GPT 5.2 on SkillsBench, exceeding the strongest closed-world baseline by 8.9 and 8.8 points, respectively. Its virtual verifier is nontrivial despite never observing hidden tests: on 84 analyzed cases it reaches 56.9% precision, 80.5% recall, 60.7% agreement, and covers 88.9% of ground-truth test intents. Transfer experiments show the same skills, produced by Opus 4.6, generalize to weaker models without adaptation, while ablations show open-world retrieval and the virtual verifier each contribute about 6 percentage points and are largely complementary; too many refinement rounds overfit, with performance peaking at 3 iterations.
Original abstract
Self-evolving agents requires adaptation after deployment, but existing approaches assume a usable learning loop, such as curated skills, successful trajectories, or verifier signals. Real open-world deployments may provide none of these, offering only a task prompt. In this work, we study open-world self-evolution, where an agent must build both its skills and its own verification signals from scratch, using open-world resources but no target-task supervision. We propose OpenSkill, a framework that bootstraps this loop: it acquires grounded knowledge and verification anchors from documentation, repositories, and the web, synthesizes them into transferable skills, and refines those skills against self-built virtual tasks grounded in the anchors rather than in target answers. The open world thus supplies both the knowledge to be learned and a supervision-independent practice environment, with target-task supervision reserved for final evaluation. Across three benchmarks and two target agents, OpenSkill attains the best automated pass rate while satisfying the no-supervision constraint. Analysis shows its skills transfer across models without model-specific adaptation, and its self-built verifier aligns with ground-truth outcomes despite never accessing them.
Read the original paperMore in AI Agents
Browse all 56 papers →LEGO-Anything: Coding Agents for 3D Scene Reconstruction
Xirui Li, Peng Shi, Mingwen Dong, Sheng Zhang, Zhuoyan Xu, Dongkyu Lee, Shuaichen Chang, Yi Xiang, Lin Pan, Jiarong Jiang
LEGO-Anything turns images into editable Blender programs through iterative coding agents, offering a promising but still imperfect route to reconstructable 3D worlds.
MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution
Prithwish Jana, Mononito Goswami, Hao Liu, Xinyu Li, Langlin Huang, Zhehui Huang, Zhishen Huang, Patrick Blöbaum, Anoop Deoras, Purak Jain, Nikos Kanakaris, Sahika Genc
MILO uses teams of evolving AI agents to automatically discover better harnesses for long-horizon problem-solving systems.
Self-Organizing Agent Teams Learn to Reason Together
Aneesh Pappu, Mirac Suzgun, Yongchan Kwon, Federico Bianchi, Batu El, Mykel J. Kochenderfer, Hancheng Cao, James Zou
This work trains AI agents to discover how to divide labor, challenge ideas, and combine reasoning so that teams can solve problems no individual agent could solve alone.