MMG2Skill: Can Agents Distill In-the-Wild Guides into Self-Evolving Skills?
AuthorsXinyu Che, Junqi Xiong, Yunfei Ge, Xinping Lei, Shihao Li, Hang Yan, Han Li, Yuanxing Zhang, Zhiqi Bai, Jinhua Hao, Ming Sun, Han Li, Jiaheng Liu
Resources
This paper teaches AI agents to turn messy human web guides into editable skills they can keep improving from their own experience.
Key results
Total success-inferable tasks in MMG2Skill-Bench
Lower end of macro-average improvement across backbones
Upper end of macro-average improvement across backbones
Largest single-cell improvement, on Game
Maximum attempt savings from analyzer-based early stopping
What the paper found
This paper introduces MMG2Skill-Bench, the first benchmark for guide-to-skill learning from in-the-wild multimodal web guides, spanning 130 success-inferable tasks across GUI control with OSWorld, open-ended Minecraft gameplay with MineStudio/OpenHA, and strategic card play with RLCard. It also proposes MMG2Skill, a closed-loop framework that converts HTML-plus-image tutorials into editable SKILL.md procedures, conditions a fixed VLM agent on those skills during execution, and revises them from trajectory-level root-cause analysis without using benchmark scores. Across six VLM backbones, including Claude-Opus-4.6, GPT-5.5, Claude-Sonnet-4.6, Kimi-K2.6, Gemini-3.1-Pro-Preview, and Qwen3.6-Plus, the method improves every model–domain cell over vanilla agents, with macro-average gains of +12.8 to +25.3 percentage points and a largest single-cell gain of +33.33 on Game. The ablation shows that raw guide prompting can be flat or harmful, while structured skill extraction and trajectory-driven revision are both necessary, especially in Game and Strategy where revision contributes over 90% of the total gain. For deployment, analyzer-based early stopping is calibrated on success-inferable tasks and saves 25.44% to 52.92% of attempts, while avoiding late-stage regressions that make full-run performance worse than early-stop on Game and Strategy.
Original abstract
Abundant procedural knowledge on the Web holds great potential for helping agents solve long-horizon tasks. However, such knowledge is often multimodal, heterogeneous, noisy, and implicitly assumes human executors, making it difficult to use directly as the skills required by agents. To bridge the gap between human-oriented guides and agent-executable skills, we formalize this problem as guide-to-skill learning: converting in-the-wild guides into executable skills and continuously improving them from trajectories observable to the agent. To evaluate the capability of existing agents on this task, we introduce MMG2Skill-Bench, the first benchmark designed for this problem. We further propose MMG2Skill, a closed-loop framework that compiles guides into editable skills, conditions a fixed vision-language model (VLM) agent on these skills during execution, and revises the skills from trajectory-level root-cause feedback without using benchmark scores. Across GUI control, open-ended gameplay, and strategic card play with six VLM backbones, MMG2Skill consistently outperforms vanilla baseline agents in every model-domain setting, achieving macro-average gains of +12.8 to +25.3 percentage points across backbones. Ablation studies show that directly prompting agents with raw guides can degrade performance, while both structured skill construction and trajectory-driven revision are necessary for the observed improvements. On success-inferable tasks, analyzer-based early stopping further prevents late-stage performance regressions and saves 25%-53% of attempts when the success signal is properly calibrated.
Read the original paperMore in AI Agents
Browse all 56 papers →LEGO-Anything: Coding Agents for 3D Scene Reconstruction
Xirui Li, Peng Shi, Mingwen Dong, Sheng Zhang, Zhuoyan Xu, Dongkyu Lee, Shuaichen Chang, Yi Xiang, Lin Pan, Jiarong Jiang
LEGO-Anything turns images into editable Blender programs through iterative coding agents, offering a promising but still imperfect route to reconstructable 3D worlds.
MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution
Prithwish Jana, Mononito Goswami, Hao Liu, Xinyu Li, Langlin Huang, Zhehui Huang, Zhishen Huang, Patrick Blöbaum, Anoop Deoras, Purak Jain, Nikos Kanakaris, Sahika Genc
MILO uses teams of evolving AI agents to automatically discover better harnesses for long-horizon problem-solving systems.
Self-Organizing Agent Teams Learn to Reason Together
Aneesh Pappu, Mirac Suzgun, Yongchan Kwon, Federico Bianchi, Batu El, Mykel J. Kochenderfer, Hancheng Cao, James Zou
This work trains AI agents to discover how to divide labor, challenge ideas, and combine reasoning so that teams can solve problems no individual agent could solve alone.