NTH

RESOURCE2SKILL: Distilling Executable Agent Skills from Human-Created Multimodal Resources

AuthorsYijia Fan, Zonglin Di, Zimo Wen, Yifan Yang, Mingxi Cheng, Qi Dai, Bei Liu, Kai Qiu, Yue Dong, Ji Li, Chong Luo

July 21, 2026 2 min read
Watch on YouTube
The one-line take

RESOURCE2SKILL turns videos, code, articles, and other human resources into reusable multimodal skills that make software agents more capable.

Key results

+11.9%
Average improvement over no-skill agents

Average overall-score lift across seven authoring domains.

26
Model-domain wins

Resource2Skill outperformed strong harness baselines in 26 of 28 aggregate cells.

+21.6%
Novel-task online acquisition gain

Adding online skills raised Tnovel performance from 41.2% to 62.8%.

68.9%
MetaBrowse selection score

Hierarchy-then-language-model selection, compared with 57.3% for no-skill agents.

85.5%
Human non-tied win rate

Skill-enabled artifacts preferred in blinded human comparisons excluding ties.

What the paper found

Resource2Skill, developed by Microsoft Research with collaborators at UC Santa Cruz and Shanghai Jiao Tong University, converts human-created tutorial videos, code repositories, articles, and reference artifacts into executable procedural knowledge for software agents. Its central artifact is a hierarchical multimodal Skill Wiki: each entry combines structured instructions, visual examples, adaptable or executable code, metadata, and provenance. At runtime, MetaBrowse uses BM25 to narrow the taxonomy, then a language model selects complementary skills that execute through MCP tools; the same distillation operator can search for missing capabilities online. Across Web, Excel, Reaper, PowerPoint, Blender, CAD, and Unreal Engine 5, Resource2Skill raises average artifact quality by +11.9 percentage points over agents without skills and wins in 26 of 28 model-domain aggregate cells, outperforming Anthropic’s Claude Code and OpenAI’s Codex harness baselines. On novel tasks, an offline pool of 891 skills augmented with 100 online skills improves performance by +21.6 percentage points, from 41.2% to 62.8%. Ablations show that the full multimodal wiki and MetaBrowse selection outperform text-only, embedding, and no-skill alternatives; the hierarchy-then-language-model selector reaches 68.9%, versus 57.3% without skills. A blinded human study also favors the skill-enabled system in 85.5% of non-tied comparisons, while failure cases reveal unresolved parameter binding and overly literal skill composition as remaining limitations.

Original abstract

Skills are a useful abstraction for software agents, turning human and agent experience into reusable procedural knowledge. Yet existing skill libraries are mostly hand-written, text-centric, or derived from agent traces, leaving tutorial videos and other multimodal human resources largely underused. We present RESOURCE2SKILL, a framework that distills multimodal resources, including tutorial videos, repositories, articles, and reference artifacts, into executable skills for software agents. RESOURCE2SKILL organizes these skills as a hierarchical multimodal Skill Wiki, where each entry combines structured text, code, visual examples, metadata, and provenance. This design preserves complementary signals from different resources: videos capture temporal operations and visual effects, code captures executable tool patterns, and articles or artifacts provide conceptual and stylistic grounding. At inference time, agents retrieve and compose relevant skills from the wiki; when coverage is insufficient, the same construction operator can acquire new skills online. Across seven practical authoring domains, RESOURCE2SKILL improves average overall score by +11.9 percentage points over no-skill agents and outperforms strong harness baselines in 26 of 28 main-aggregate model-domain cells. Ablations confirm the value of multimodal skill format, hierarchical organization, source diversity, selection strategy, and online acquisition.

Read the original paper

More in AI Agents

Browse all 56 papers →
01Agent

LEGO-Anything: Coding Agents for 3D Scene Reconstruction

Xirui Li, Peng Shi, Mingwen Dong, Sheng Zhang, Zhuoyan Xu, Dongkyu Lee, Shuaichen Chang, Yi Xiang, Lin Pan, Jiarong Jiang

LEGO-Anything turns images into editable Blender programs through iterative coding agents, offering a promising but still imperfect route to reconstructable 3D worlds.

Read analysis
02Agent

MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution

Prithwish Jana, Mononito Goswami, Hao Liu, Xinyu Li, Langlin Huang, Zhehui Huang, Zhishen Huang, Patrick Blöbaum, Anoop Deoras, Purak Jain, Nikos Kanakaris, Sahika Genc

MILO uses teams of evolving AI agents to automatically discover better harnesses for long-horizon problem-solving systems.

Read analysis
03Agent

Self-Organizing Agent Teams Learn to Reason Together

Aneesh Pappu, Mirac Suzgun, Yongchan Kwon, Federico Bianchi, Batu El, Mykel J. Kochenderfer, Hancheng Cao, James Zou

This work trains AI agents to discover how to divide labor, challenge ideas, and combine reasoning so that teams can solve problems no individual agent could solve alone.

Read analysis