Local Support Learning
AuthorsAssaf Ben-Kish, Akarsh Kumar, James Glass, Raja Giryes
AffiliationsMIT CSAIL · Tel Aviv University
Resources
Local Support Learning helps large language models learn new skills without overwriting what they already know by activating updates only where they are locally needed.
Key results
LSL retained pretrained and fine-tuned capabilities on Qwen2.5-7B-Instruct.
The GMM gate improved pretrained-capability retention from 76.6% to 96.6%.
Temporal smoothing raised retention from 96.6% to 98.8%.
Both positive and negative GMM gates required 13.8MB of additional memory.
Each LSL adapter added 86ms per forward pass.
Fitting the GMMs increased total training time by 1.64x on 1M tokens.
What the paper found
Local Support Learning, or LSL, addresses catastrophic forgetting in large language models by treating each weight matrix as a mapping over an activation space: conventional gradient updates modify that mapping for every input, including inputs from earlier training phases. LSL combines a standard LoRA-style adapter with a per-matrix Gaussian Mixture Model gate, applying the adapter only when a token’s activation resembles the current phase’s distribution. A second GMM provides an input-dependent negative reference, while temporal smoothing stabilizes token-level routing without requiring access to prior data. On Qwen2.5-7B-Instruct, LSL was evaluated with chemistry, English-to-Igbo translation, and cybersecurity fine-tuning, while measuring retention on GSM8K, HumanEval, and IFEval. The GMM gate raised pretrained-capability retention from 76.6% to 96.6%, and smoothing increased it to 98.8%, while preserving full downstream learning across sequential phases. The method scales from 1.5B to 7B models, adds 13.8MB of GMM memory, incurs 86ms of inference latency per adapter forward pass, and increases training time by 1.64x on 1M tokens relative to LoRA. Unlike LoRA, OP-LoRA, and Learning without Forgetting, LSL largely decouples new-task learning from retention, although its adapters cannot be merged into the base weights and inference cost grows with the number of phases.
Original abstract
We explore catastrophic forgetting in the context of large pre-trained models. By considering forgetting as a geometric problem in the input space of each weight matrix, we uncover a natural retention objective under which updates produced by gradient-based optimizers are suboptimal. Following this observation, we propose Local Support Learning (LSL), a general-purpose framework that augments gradient-based training for retention of prior capabilities without access to prior data. During a new learning phase, LSL pairs two components with distinct roles: a standard weight adapter, trained as usual to minimize the loss, and a gating function that enables the adapter only on input activations from its own training distribution, making the update local to that distribution. The key challenge is that this gate must route data from all learning phases while training only on data from the current one. We address this with a gate based on a Gaussian Mixture Model (GMM), whose likelihood decays rapidly away from its training data, giving it a natural tendency to stay closed on data from prior phases. We show that this post-training approach can resolve forgetting in LLMs of up to 7 billion parameters, retaining both pretrained and finetuned capabilities across multiple training phases, while being efficient in memory and compute, robust to hyperparameter choice, and showing scaling potential.
Read the original paperMore in Continual Learning
Browse all 24 papers →ASCENT: Online Test-Time Training of Long-Horizon Agents via Self-Distillation of Verified Experience
Haodong Lu, Dong Gong
ASCENT lets deployed LLM agents learn from verified successes on the fly by converting hindsight about their own trajectories into lasting weight updates.
From Knowledge Access to Source Learning: Developing Source-Specific Competence
Lucheng Fu, Kejing Xia, Yiyang Wang, Yiqiao Jin, Jinjin He, Xiyuan Yang, Haoxin Liu, Ye Yu, Haibo Jin, Yijia Xiao, Wenke Lee, B. Aditya Prakash, Haohan Wang
SourceLearn helps LLM agents progressively build reusable expertise about trusted information sources instead of repeatedly starting from scratch.
ACLArena: Agent Continue Learning in Multi-stage Post-training
Haixin Wang, Xiaoxuan Wang, Junkai Zhang, Han Zhang, Renliang Sun, Alexander K Taylor, Yidan Shi, Haoran Deng, Chenguang Wang, Jason Cong, Yizhou Sun, Wei Wang
ACLArena studies how agents can learn new skills over multiple training stages without forgetting old ones, proposing replay and specialized LoRA experts as a practical solution.