NTH

Harness Continual Learning: Continual Adaptation Beyond Model Parameters

AuthorsBorui Kang, Jinrui Gu, Junhan Lv, Wenbin Li, Lei Wang, Yang Gao

August 21, 2026 2 min read
Watch on YouTube
The one-line take

This paper proposes letting frozen AI models keep learning through evolving prompts, memories, tools, and routing rules while guarding against the loss of earlier skills.

Key results

62.98%
ALFWorld Plasticity-HCL final average

Best final average across six ALFWorld task categories.

64.70%
Textual reasoning Plasticity-HCL

Final average across MuSiQue, ProofWriter, GSM8K, and HotpotQA.

45.50%
Textual zero-shot baseline

DeepSeek-V4-Flash evaluated without sequential harness updates.

68.92%
Multimodal Stability-HCL

Final average across COCO detection, COCO captioning, RefCOCO, and VQAv2.

0.22
Multimodal forgetting

Average forgetting achieved by Stability-HCL.

63.46%
Retention sweep peak

Highest final average obtained with historical-loss tolerance b=1.

What the paper found

The paper introduces Harness Continual Learning, or HCL, which shifts continual adaptation from model parameters to the external harness surrounding a frozen foundation model such as DeepSeek-V4-Flash or Qwen3.6-27B. The evolving state has four jointly versioned components: a Task Interface, Experience Memory, Capability Map, and Adaptive Router, covering prompts, raw and abstract memories, reusable skills, tools, and workflow policies. HCL identifies harness-level forgetting, where a memory, skill, or routing update breaks previously reliable answers, tool calls, or action trajectories without changing the model. Its guarded evolution separates candidate generation from deployment: a Continual Optimizer proposes revisions using execution feedback, while a Continual Evaluator commits them only after current-task improvement, historical-anchor retention, and validity checks. On ALFWorld, Plasticity-HCL reaches a 62.98% final average, versus 55.56% for retrieval-augmented generation, while Stability-HCL reaches 61.74% with lower forgetting. In textual reasoning across MuSiQue, ProofWriter, GSM8K, and HotpotQA, Plasticity-HCL achieves 64.70%, compared with 45.50% zero-shot, with 0.07 average forgetting. On multimodal COCO detection, COCO captioning, RefCOCO grounding, and VQAv2, Stability-HCL reaches 68.92% with 0.22 forgetting. A retention sweep shows that moderate plasticity is superior: tolerance b=1 produces 63.46%, while unrestricted b=∞ falls to 60.13% and forgetting rises to 3.45. The results show that reliable agent improvement requires evolving and validating the harness, not merely adding memory or updating model weights.

Original abstract

Continual learning has largely been model-centric, treating model parameters as the state that changes with sequential experience. Modern agents can also adapt through a harness of prompts, memories, tools, skills, and routing rules. Because these contents jointly shape later execution, a harness update can disrupt previously reliable behavior even when the model is frozen. This raises a new question: how can an agent continually improve its state outside the model while retaining behavior acquired earlier? We formulate Harness Continual Learning (HCL), a new continual learning paradigm in which the harness evolves around a frozen foundation model, and define the resulting loss of earlier behavior as harness-level forgetting. We instantiate HCL with four execution-facing components: the Task Interface, Experience Memory, Capability Map, and Adaptive Router. We further introduce guarded harness evolution to separate update generation from state commitment. A Continual Optimizer proposes candidate harnesses from post-execution feedback, and a Continual Evaluator commits the resulting candidate harness only after checking current improvement, historical retention, and validity. Experiments on textual reasoning, multimodal perception, and open-world interaction demonstrate capability accumulation and failure recovery, with relative gains exceeding 10% over corresponding baselines in multiple settings. Component ablations assess the contribution of each harness component, while controlled retention sweeps reveal measurable harness-level forgetting and show that the stability--plasticity trade-off can be explicitly adjusted.

Read the original paper

More in Continual Learning

Browse all 24 papers →
02Continual Learning

From Knowledge Access to Source Learning: Developing Source-Specific Competence

Lucheng Fu, Kejing Xia, Yiyang Wang, Yiqiao Jin, Jinjin He, Xiyuan Yang, Haoxin Liu, Ye Yu, Haibo Jin, Yijia Xiao, Wenke Lee, B. Aditya Prakash, Haohan Wang

SourceLearn helps LLM agents progressively build reusable expertise about trusted information sources instead of repeatedly starting from scratch.

Read analysis
03Continual Learning

Local Support Learning

Assaf Ben-Kish, Akarsh Kumar, James Glass, Raja Giryes

Local Support Learning helps large language models learn new skills without overwriting what they already know by activating updates only where they are locally needed.

Read analysis