ContinualSkillBench: Can LLM Agents Truly Evolve Their Capabilities?
AuthorsTianyi Guan, Yiding Wang, Haotong Yang, Siyuan Cao, Shirui Liu, Yi Hu, Jiaqi Li, Muhan Zhang
Resources
ContinualSkillBench tests whether LLM agents truly learn reusable skills over time, finding that they often adapt to context without reliably consolidating transferable abilities.
Key results
Each domain contains an ordered stream of 100 interconnected subtasks.
Relative improvement in normalized reward over independent execution.
Average normalized reward for pure in-context learning, versus 0.602 with explicit skill maintenance.
Total generated skills across the five domains.
Average quality score for its generated skill pool, compared with 5.68 for GPT-4o.
What the paper found
ContinualSkillBench tests whether LLM agents can convert experience into reusable capabilities rather than merely benefit from longer context. It organizes five domains—Healthcare, Law, Mathematics, Finance, and Office—into ordered streams of 100 interconnected subtasks, using feedback after each task and a three-turn instruction, execution, and reflection protocol. Across GPT-4o, GPT-5.3-Codex, and Claude 4.7 Opus, sequential execution improves normalized reward in 14 of 15 model–domain combinations, producing a 16.9% aggregate relative gain over independent execution. However, an ablation with GPT-5.3-Codex shows that pure in-context learning reaches 0.605 average normalized reward, compared with 0.602 for explicit skill maintenance, indicating that retained context and feedback explain much of the improvement. Explicit skills still help with reusable procedures, exact outputs, and programmatic tasks. Skill repositories reveal a capability gap: GPT-4o generates 384 skills with an average quality score of 5.68, while the stronger GPT-5.3-Codex maintains 205 skills scoring 7.94 on average, suggesting better consolidation and reuse. Built on Harbor with Codex CLI and Claude Code harnesses, the benchmark concludes that current agents can adapt continually but still struggle to compress experience into robust, transferable abstractions.
Original abstract
Modern agent frameworks equip large language models with external skill libraries to solve complex tasks. However, it remains unclear whether these systems can effectively evolve their skills and whether the resulting skills improve task-solving capabilities. To bridge this gap, we introduce ContinualSkillBench, a dynamic evaluation framework for in-context continual skill learning. It covers five representative domains, each containing 100 interconnected subtasks ordered by increasing difficulty and opportunities for cross-task skill reuse. Our experiments show that sequential execution generally improves performance, but the gains vary substantially across models and domains. Moreover, in-context learning performs comparably to explicit skill maintenance on average, suggesting that much of the improvement arises from adaptation to prior context and feedback rather than reusable skill abstraction alone. Explicit skills nevertheless provide selective benefits for tasks requiring reusable procedures or precise outputs. We further find that less capable models tend to accumulate larger, more fragmented collections of task-specific skills. These findings show that current in-context skill evolution mechanisms can support continual adaptation, but still struggle to consistently consolidate experience into robust and transferable skills.
Read the original paperMore in Continual Learning
Browse all 24 papers →ASCENT: Online Test-Time Training of Long-Horizon Agents via Self-Distillation of Verified Experience
Haodong Lu, Dong Gong
ASCENT lets deployed LLM agents learn from verified successes on the fly by converting hindsight about their own trajectories into lasting weight updates.
From Knowledge Access to Source Learning: Developing Source-Specific Competence
Lucheng Fu, Kejing Xia, Yiyang Wang, Yiqiao Jin, Jinjin He, Xiyuan Yang, Haoxin Liu, Ye Yu, Haibo Jin, Yijia Xiao, Wenke Lee, B. Aditya Prakash, Haohan Wang
SourceLearn helps LLM agents progressively build reusable expertise about trusted information sources instead of repeatedly starting from scratch.
Local Support Learning
Assaf Ben-Kish, Akarsh Kumar, James Glass, Raja Giryes
Local Support Learning helps large language models learn new skills without overwriting what they already know by activating updates only where they are locally needed.