NTH

Local Support Learning

AuthorsAssaf Ben-Kish, Akarsh Kumar, James Glass, Raja Giryes

AffiliationsMIT CSAIL · Tel Aviv University

October 6, 2026 2 min read
Watch on YouTube
The one-line take

Local Support Learning helps large language models learn new skills without overwriting what they already know by activating updates only where they are locally needed.

Key results

7B
Largest evaluated model

LSL retained pretrained and fine-tuned capabilities on Qwen2.5-7B-Instruct.

96.6%
GMM retention

The GMM gate improved pretrained-capability retention from 76.6% to 96.6%.

98.8%
Smoothed retention

Temporal smoothing raised retention from 96.6% to 98.8%.

13.8MB
GMM memory overhead

Both positive and negative GMM gates required 13.8MB of additional memory.

86ms
Inference latency

Each LSL adapter added 86ms per forward pass.

1.64x
Training-time multiplier

Fitting the GMMs increased total training time by 1.64x on 1M tokens.

What the paper found

Local Support Learning, or LSL, addresses catastrophic forgetting in large language models by treating each weight matrix as a mapping over an activation space: conventional gradient updates modify that mapping for every input, including inputs from earlier training phases. LSL combines a standard LoRA-style adapter with a per-matrix Gaussian Mixture Model gate, applying the adapter only when a token’s activation resembles the current phase’s distribution. A second GMM provides an input-dependent negative reference, while temporal smoothing stabilizes token-level routing without requiring access to prior data. On Qwen2.5-7B-Instruct, LSL was evaluated with chemistry, English-to-Igbo translation, and cybersecurity fine-tuning, while measuring retention on GSM8K, HumanEval, and IFEval. The GMM gate raised pretrained-capability retention from 76.6% to 96.6%, and smoothing increased it to 98.8%, while preserving full downstream learning across sequential phases. The method scales from 1.5B to 7B models, adds 13.8MB of GMM memory, incurs 86ms of inference latency per adapter forward pass, and increases training time by 1.64x on 1M tokens relative to LoRA. Unlike LoRA, OP-LoRA, and Learning without Forgetting, LSL largely decouples new-task learning from retention, although its adapters cannot be merged into the base weights and inference cost grows with the number of phases.

Original abstract

We explore catastrophic forgetting in the context of large pre-trained models. By considering forgetting as a geometric problem in the input space of each weight matrix, we uncover a natural retention objective under which updates produced by gradient-based optimizers are suboptimal. Following this observation, we propose Local Support Learning (LSL), a general-purpose framework that augments gradient-based training for retention of prior capabilities without access to prior data. During a new learning phase, LSL pairs two components with distinct roles: a standard weight adapter, trained as usual to minimize the loss, and a gating function that enables the adapter only on input activations from its own training distribution, making the update local to that distribution. The key challenge is that this gate must route data from all learning phases while training only on data from the current one. We address this with a gate based on a Gaussian Mixture Model (GMM), whose likelihood decays rapidly away from its training data, giving it a natural tendency to stay closed on data from prior phases. We show that this post-training approach can resolve forgetting in LLMs of up to 7 billion parameters, retaining both pretrained and finetuned capabilities across multiple training phases, while being efficient in memory and compute, robust to hyperparameter choice, and showing scaling potential.

Read the original paper

More in Continual Learning

Browse all 24 papers →
02Continual Learning

From Knowledge Access to Source Learning: Developing Source-Specific Competence

Lucheng Fu, Kejing Xia, Yiyang Wang, Yiqiao Jin, Jinjin He, Xiyuan Yang, Haoxin Liu, Ye Yu, Haibo Jin, Yijia Xiao, Wenke Lee, B. Aditya Prakash, Haohan Wang

SourceLearn helps LLM agents progressively build reusable expertise about trusted information sources instead of repeatedly starting from scratch.

Read analysis
03Continual Learning

ACLArena: Agent Continue Learning in Multi-stage Post-training

Haixin Wang, Xiaoxuan Wang, Junkai Zhang, Han Zhang, Renliang Sun, Alexander K Taylor, Yidan Shi, Haoran Deng, Chenguang Wang, Jason Cong, Yizhou Sun, Wei Wang

ACLArena studies how agents can learn new skills over multiple training stages without forgetting old ones, proposing replay and specialized LoRA experts as a practical solution.

Read analysis