NTH

ACLArena: Agent Continue Learning in Multi-stage Post-training

AuthorsHaixin Wang, Xiaoxuan Wang, Junkai Zhang, Han Zhang, Renliang Sun, Alexander K Taylor, Yidan Shi, Haoran Deng, Chenguang Wang, Jason Cong, Yizhou Sun, Wei Wang

AffiliationsStage 2 ... Stage K Teacher Β· Teacher 2 ... Teacher K πŸ” (Ο€0, D1) πŸ”₯ (Ο€1, D2) πŸ”₯ ... (Ο€K-1, DK) πŸ”₯ ... Generate Model Ο€1 Model Ο€2 ... Final Model Dist Β· University of California, Santa Cruz

September 24, 2026 2 min read
Watch on YouTube
The one-line take

ACLArena studies how agents can learn new skills over multiple training stages without forgetting old ones, proposing replay and specialized LoRA experts as a practical solution.

Key results

17K
Math training dataset

DAPO-Math-17K is used for the mathematical reasoning stage.

6.04
Sequential AIME26 after e-commerce

AIME26 declined from 25.83 after math training to 6.04 after e-commerce training.

49.7
MLE NQ score

Mixture of Low-Rank Experts reaches 49.7 on the NQ search benchmark.

38.6
MLE multi-hop search

MLE achieves 38.6 on out-of-domain multi-hop search.

42.4
MLE GPQA score

MLE reaches 42.4 on GPQA.

85.0
MLE IF-Eval score

MLE achieves 85.0 on the in-domain instruction-following benchmark.

What the paper found

ACLArena studies why agents forget earlier skills when post-training proceeds through multiple stages: mathematical reasoning, search, e-commerce tool use, and instruction following. Using Qwen3-8B-Base, the benchmark compares sequential training with multi-teacher mixed on-policy distillation, self-distilled fine-tuning, and model merging, connecting the analysis to consolidation recipes used in Qwen3, DeepSeek-V4, and Microsoft’s MAI-Thinking-1. Sequential updates move parameters in only partially aligned task directions; token analysis shows that forgetting concentrates at high-entropy decision points rather than uniformly across generations. The effect is severe: AIME26 falls from 25.83 after math training to 6.04 after e-commerce training, while search performance also regresses. The proposed solution, Mixture of Low-Rank Experts, first uses filtered expert trajectories for SDFT, then freezes the shared backbone and trains a separate LoRA adapter with reinforcement learning for each domain, using environment context for routing. Across the four-stage pipeline, MLE reaches 49.7 on NQ, 38.6 on multi-hop search, and 42.4 on GPQA, while achieving 85.0 on IF-Eval. The study also uses DAPO-Math-17K and finds that SFT establishes valid tool-use behavior through larger parameter updates, whereas RL and mixed OPD make smaller, policy-local refinements. Overall, ACLArena argues that continual agent learning should share transferable behavior but isolate conflicting task-specific adaptations instead of forcing every capability into one parameter space.

Original abstract

Building general-purpose agents for industrial deployment requires integrating multiple capabilities, each typically acquired at a distinct stage of training. Yet there is currently no well-established recipe for Agent Continual Learning (ACL), with little understanding of the trade-offs among existing integration paradigms. To address this gap, we introduce ACLArena, a framework for comprehensively studying, analyzing, and evaluating ACL. We first build a sequential training pipeline and conduct an in-depth analysis that explains the mechanisms of forgetting and generalization from two complementary perspectives, the model level and the token level. Guided by these analyses, we systematically compare multi-teacher on-policy distillation, self-distilled fine-tuning, and model merging to assess their ability to recover previously learned capabilities while preserving newly acquired ones. Through extensive experiments, we develop a detailed understanding of how capabilities transfer across stages. Finally, we propose a new ACL recipe that combines offline replay over high-quality trajectories with a routed network of multiple LoRA experts each specialized via RL, substantially improving the agent's ability to learn across multiple domains. Comprehensive experiments on four reasoning and agentic tasks, evaluated under both in-domain and out-of-domain settings, demonstrate the value of our analysis and the effectiveness of our approach.

Read the original paper

More in Continual Learning

Browse all 24 papers β†’
02Continual Learning

From Knowledge Access to Source Learning: Developing Source-Specific Competence

Lucheng Fu, Kejing Xia, Yiyang Wang, Yiqiao Jin, Jinjin He, Xiyuan Yang, Haoxin Liu, Ye Yu, Haibo Jin, Yijia Xiao, Wenke Lee, B. Aditya Prakash, Haohan Wang

SourceLearn helps LLM agents progressively build reusable expertise about trusted information sources instead of repeatedly starting from scratch.

Read analysis
03Continual Learning

Local Support Learning

Assaf Ben-Kish, Akarsh Kumar, James Glass, Raja Giryes

Local Support Learning helps large language models learn new skills without overwriting what they already know by activating updates only where they are locally needed.

Read analysis