SkillGate: Training In-Policy Skill Selection in Long-Horizon Agents
AuthorsQingyao Li, Wenxiang Jiao, Shuai Shao, Kangning Zhang, Yuan Lu, Yi Guo, Weiwen Liu, Weinan Zhang, Yong Yu
Resources
SkillGate helps long-horizon AI agents choose the right procedural skill by giving skill-selection decisions their own dedicated learning signal.
Key results
On-policy trajectories used to diagnose selector credit starvation.
Share of trajectory loss weight carried by skill-naming tokens.
Percentage-point success advantage associated with reading the correct skill.
Overall success of the Qwen3.5-9B policy across five benchmarks.
Controlled baseline using the same training budget.
SkillGate misleading exposure, versus 69.6% for outcome-only SkillRL.
What the paper found
SkillGate addresses a failure in long-horizon agents that select procedural skills during execution: standard outcome-rewarded GRPO broadcasts one sequence-level advantage across thousands of tokens, starving the few tokens that name the chosen skill. An audit of 12,800 on-policy trajectories found that skill-identity tokens carried a median 0.14% of trajectory loss weight, while choosing the correct skill was associated with an 11.2 percentage-point success advantage; longer trajectories also made the inherited credit increasingly wrong-signed. SkillGate separates credit into two disjoint channels: outcome advantage trains execution tokens only, with the entire read call removed, while an action-local selector advantage trains exactly the skill-name tokens and is positive only for a single read of the oracle skill. On five benchmarks—Claw-Eval, SkillsBench, SETA, SWE, and Terminal-Bench 2.0—with 16 candidates per slate, SkillGate raises the Qwen3.5-9B policy’s trial success from 47.0% with outcome-only SkillRL to 53.2%, outperforming larger references including Qwen3.5-397B-A17B and DeepSeek-V3.2. It also shifts behavior from indiscriminate retrieval toward selective access: misleading-skill exposure falls from 69.6% to 21.8%, while the agent reads fewer skills. The results suggest that direct token-level training of in-policy selection can matter more than simply scaling the model or adding an external router.
Original abstract
Agent frameworks increasingly package procedural knowledge as skills: instruction files an agent reads on demand, while public libraries now hold thousands of them. Which skill to read has thus become a decision the policy itself makes in the middle of an episode, yet no existing signal trains it. We show that the default remedy, outcome-rewarded RL over the candidate slate, cannot teach it, for a structural reason we identify and name selector credit starvation: under a broadcast, sequence-level advantage, the few tokens that name the chosen skill carry a vanishing share of the loss, and the credit they inherit is increasingly wrong-signed as trajectories lengthen. A correct choice is punished whenever the execution after it fails, even though the choice itself is among the most valuable decisions in the trajectory. Auditing a completed run's own training artifacts confirms all three properties, each worsening monotonically with horizon. SkillGate removes the failure by construction: it partitions the token support into two disjoint credit channels, outcome credit reaching only execution tokens, and a separate action-local advantage reaching exactly the skill-naming tokens, positive only when a trajectory's single read is the correct one. On five agentic benchmarks under a 16-candidate slate, SkillGate lifts a 9B policy from 40.8% to 53.2% trial success, well ahead of the identical budget spent on outcome reward alone, while cutting exposure to misleading candidates by two thirds and reading fewer skills.
Read the original paperMore in AI Agents
Browse all 56 papers →LEGO-Anything: Coding Agents for 3D Scene Reconstruction
Xirui Li, Peng Shi, Mingwen Dong, Sheng Zhang, Zhuoyan Xu, Dongkyu Lee, Shuaichen Chang, Yi Xiang, Lin Pan, Jiarong Jiang
LEGO-Anything turns images into editable Blender programs through iterative coding agents, offering a promising but still imperfect route to reconstructable 3D worlds.
MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution
Prithwish Jana, Mononito Goswami, Hao Liu, Xinyu Li, Langlin Huang, Zhehui Huang, Zhishen Huang, Patrick Blöbaum, Anoop Deoras, Purak Jain, Nikos Kanakaris, Sahika Genc
MILO uses teams of evolving AI agents to automatically discover better harnesses for long-horizon problem-solving systems.
Self-Organizing Agent Teams Learn to Reason Together
Aneesh Pappu, Mirac Suzgun, Yongchan Kwon, Federico Bianchi, Batu El, Mykel J. Kochenderfer, Hancheng Cao, James Zou
This work trains AI agents to discover how to divide labor, challenge ideas, and combine reasoning so that teams can solve problems no individual agent could solve alone.