NTH

MMG2Skill: Can Agents Distill In-the-Wild Guides into Self-Evolving Skills?

AuthorsXinyu Che, Junqi Xiong, Yunfei Ge, Xinping Lei, Shihao Li, Hang Yan, Han Li, Yuanxing Zhang, Zhiqi Bai, Jinhua Hao, Ming Sun, Han Li, Jiaheng Liu

June 15, 2026 2 min read
Watch on YouTube
The one-line take

This paper teaches AI agents to turn messy human web guides into editable skills they can keep improving from their own experience.

Key results

130
benchmark_tasks

Total success-inferable tasks in MMG2Skill-Bench

12.8
macro_avg_gain_low

Lower end of macro-average improvement across backbones

25.3
macro_avg_gain_high

Upper end of macro-average improvement across backbones

33.33
largest_single_cell_gain

Largest single-cell improvement, on Game

52.92%
attempt_savings

Maximum attempt savings from analyzer-based early stopping

What the paper found

This paper introduces MMG2Skill-Bench, the first benchmark for guide-to-skill learning from in-the-wild multimodal web guides, spanning 130 success-inferable tasks across GUI control with OSWorld, open-ended Minecraft gameplay with MineStudio/OpenHA, and strategic card play with RLCard. It also proposes MMG2Skill, a closed-loop framework that converts HTML-plus-image tutorials into editable SKILL.md procedures, conditions a fixed VLM agent on those skills during execution, and revises them from trajectory-level root-cause analysis without using benchmark scores. Across six VLM backbones, including Claude-Opus-4.6, GPT-5.5, Claude-Sonnet-4.6, Kimi-K2.6, Gemini-3.1-Pro-Preview, and Qwen3.6-Plus, the method improves every model–domain cell over vanilla agents, with macro-average gains of +12.8 to +25.3 percentage points and a largest single-cell gain of +33.33 on Game. The ablation shows that raw guide prompting can be flat or harmful, while structured skill extraction and trajectory-driven revision are both necessary, especially in Game and Strategy where revision contributes over 90% of the total gain. For deployment, analyzer-based early stopping is calibrated on success-inferable tasks and saves 25.44% to 52.92% of attempts, while avoiding late-stage regressions that make full-run performance worse than early-stop on Game and Strategy.

Original abstract

Abundant procedural knowledge on the Web holds great potential for helping agents solve long-horizon tasks. However, such knowledge is often multimodal, heterogeneous, noisy, and implicitly assumes human executors, making it difficult to use directly as the skills required by agents. To bridge the gap between human-oriented guides and agent-executable skills, we formalize this problem as guide-to-skill learning: converting in-the-wild guides into executable skills and continuously improving them from trajectories observable to the agent. To evaluate the capability of existing agents on this task, we introduce MMG2Skill-Bench, the first benchmark designed for this problem. We further propose MMG2Skill, a closed-loop framework that compiles guides into editable skills, conditions a fixed vision-language model (VLM) agent on these skills during execution, and revises the skills from trajectory-level root-cause feedback without using benchmark scores. Across GUI control, open-ended gameplay, and strategic card play with six VLM backbones, MMG2Skill consistently outperforms vanilla baseline agents in every model-domain setting, achieving macro-average gains of +12.8 to +25.3 percentage points across backbones. Ablation studies show that directly prompting agents with raw guides can degrade performance, while both structured skill construction and trajectory-driven revision are necessary for the observed improvements. On success-inferable tasks, analyzer-based early stopping further prevents late-stage performance regressions and saves 25%-53% of attempts when the success signal is properly calibrated.

Read the original paper

More in AI Agents

Browse all 56 papers →
01Agent

LEGO-Anything: Coding Agents for 3D Scene Reconstruction

Xirui Li, Peng Shi, Mingwen Dong, Sheng Zhang, Zhuoyan Xu, Dongkyu Lee, Shuaichen Chang, Yi Xiang, Lin Pan, Jiarong Jiang

LEGO-Anything turns images into editable Blender programs through iterative coding agents, offering a promising but still imperfect route to reconstructable 3D worlds.

Read analysis
02Agent

MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution

Prithwish Jana, Mononito Goswami, Hao Liu, Xinyu Li, Langlin Huang, Zhehui Huang, Zhishen Huang, Patrick Blöbaum, Anoop Deoras, Purak Jain, Nikos Kanakaris, Sahika Genc

MILO uses teams of evolving AI agents to automatically discover better harnesses for long-horizon problem-solving systems.

Read analysis
03Agent

Self-Organizing Agent Teams Learn to Reason Together

Aneesh Pappu, Mirac Suzgun, Yongchan Kwon, Federico Bianchi, Batu El, Mykel J. Kochenderfer, Hancheng Cao, James Zou

This work trains AI agents to discover how to divide labor, challenge ideas, and combine reasoning so that teams can solve problems no individual agent could solve alone.

Read analysis