ASCENT: Online Test-Time Training of Long-Horizon Agents via Self-Distillation of Verified Experience
AuthorsHaodong Lu, Dong Gong
AffiliationsUniversity of New South Wales (UNSW Sydney)
Resources
ASCENT lets deployed LLM agents learn from verified successes on the fly by converting hindsight about their own trajectories into lasting weight updates.
Key results
Overall success rate on the 140-task ALFWorld seen stream.
Base-model success rate on the same ALFWorld stream.
Overall success rate on the 140-task ALFWorld seen stream.
Base-model success rate on the same ALFWorld stream.
Exact task success on the WebShop stream.
Exact task success on the WebShop stream.
What the paper found
ASCENT introduces Online Agentic Test-Time Training, allowing a long-horizon agent to improve its weights during deployment rather than relying on retrieved text memories or a separate training phase. For each task, the agent makes one attempt, receives a sparse verifier outcome, and updates persistent LoRA fast weights only after success. Instead of imitating its own generated tokens—which can cause policy drift and invalid actions—ASCENT uses a frozen initial copy of the open-weight Qwen3.5 model as a teacher. The teacher receives the completed successful trajectory as privileged hindsight and produces full next-token distributions; the student then matches them with forward KL at its original prefixes. An action-validity filter removes unusable turns from the teacher context while retaining the student’s training positions, creating shorter, more executable behavioral guidance. On the 140-task ALFWorld seen stream, success reaches 69.5% with Qwen3.5-4B versus 46.4% for the base, and 77.4% with Qwen3.5-9B versus 55.0%. On WebShop, strict success reaches 41.9% for Qwen3.5-4B and 41.5% for Qwen3.5-9B, while mean interaction length falls to 17.8 and 26.7 turns respectively. ASCENT also outperforms online memory and turn-local self-distillation baselines, transfers to held-out ALFWorld scenes, and improves interactive coding on AppWorld, showing that deployed agents can consolidate verified experience directly into parameters.
Original abstract
A large language model (LLM) agent solves long-horizon tasks through many reasoning-action turns, with one verification signal at termination. Deployed agents face streams of related tasks, making their trajectories a natural resource for improvement. In-context adaptation agents store reflections, memories, or skills as text, so reuse depends on retrieving the right experience and on a frozen policy executing it. We study Online Agentic Test-Time Training (OaTTT), which trains the LLM's weights on its own execution trajectories during deployment. The agent executes each task once, in one pass over the stream, and the executed trajectory with its verification result is the only learning signal for weight updates that persist across tasks. Directly imitating or reinforcing the generated tokens of this single attempt destabilizes the policy. We introduce ASCENT (Agentic Self-distillation for Cross-task EvolutioN at Test-time), which instead self-distills verified experience. A stable version of the LLM, its frozen initial copy, receives the verified trajectory as privileged information and predicts next-token distributions along it with this hindsight. Distilling them into persistent LoRA fast weights updates the agent for later tasks, without an external reference solution or stronger teacher. By further removing invalid-action turns, ASCENT distills enhanced privileged experience for more efficient execution. We characterize its population target and the limits of sparse outcome selection. Across ALFWorld, WebShop, and AppWorld at varied model scales, ASCENT improves task success and interaction efficiency as experience accumulates, outperforms online adaptation methods, and transfers to held-out scenes, showing that an agent can consolidate verified experience into its weights without a separate training phase or memory retrieval. Project page: https://artificer-ai-lab.github.io/ASCENT
Read the original paperMore in Continual Learning
Browse all 24 papers →From Knowledge Access to Source Learning: Developing Source-Specific Competence
Lucheng Fu, Kejing Xia, Yiyang Wang, Yiqiao Jin, Jinjin He, Xiyuan Yang, Haoxin Liu, Ye Yu, Haibo Jin, Yijia Xiao, Wenke Lee, B. Aditya Prakash, Haohan Wang
SourceLearn helps LLM agents progressively build reusable expertise about trusted information sources instead of repeatedly starting from scratch.
Local Support Learning
Assaf Ben-Kish, Akarsh Kumar, James Glass, Raja Giryes
Local Support Learning helps large language models learn new skills without overwriting what they already know by activating updates only where they are locally needed.
ACLArena: Agent Continue Learning in Multi-stage Post-training
Haixin Wang, Xiaoxuan Wang, Junkai Zhang, Han Zhang, Renliang Sun, Alexander K Taylor, Yidan Shi, Haoran Deng, Chenguang Wang, Jason Cong, Yizhou Sun, Wei Wang
ACLArena studies how agents can learn new skills over multiple training stages without forgetting old ones, proposing replay and specialized LoRA experts as a practical solution.