NTH

MemSFT: Mitigating Alignment Tax with an External Parametric Memory

AuthorsJiarui Wang, Xiang Shi, Jiaqi Cao, Rubin Wei, Xiquan Wang, Hao Sun, Jingzhi Wang, Zhiqi Yang, Qipeng Guo, Bowen Zhou, Zhouhan Lin

August 5, 2026 2 min read
Watch on YouTube
The one-line take

MemSFT gives LLMs specialized domain expertise through an external memory while preserving their general abilities.

Key results

42.92
BioIns Qwen3-14B MemSFT score

Average Biology-Instructions score after MemSFT.

83.62
General capability average

Qwen3-14B average across five general benchmarks after MemSFT.

0.47
OpenSWI RMSE

Best reported shallow OpenSWI RMSE for Qwen3-14B with MemSFT.

56.47
LawBench average

Qwen3-14B legal-domain score with MemSFT.

9.23
Four-backbone adaptation compute

Estimated EFLOPs for MemSFT across four Qwen3 backbones.

0.22
Relative compute cost

MemSFT compute relative to full SFT across four backbones.

What the paper found

MemSFT, from Shanghai Jiao Tong University and Shanghai AI Laboratory, addresses the alignment tax caused by domain fine-tuning: instead of updating a post-trained backbone, it trains a plug-and-play parametric memory to imitate FAISS nearest-neighbor retrieval over domain instruction data, using a combination of KL-divergence and cross-entropy losses. A lightweight token-level router then interpolates the frozen backbone’s and memory’s next-token distributions, invoking specialized knowledge selectively while preserving general reasoning and instruction following. Across Biology-Instructions, OpenSWI, and LawBench, experiments with Qwen3 models show that one domain-specific Qwen3-8B memory can transfer from Qwen3-8B through Qwen3-235B-A22B, while a LLaMA2-13B validation indicates the approach is not limited to Qwen3. On Qwen3-14B, MemSFT raises BioIns performance to 42.92 while retaining a general-benchmark average of 83.62; it achieves an OpenSWI RMSE of 0.47 and a LawBench score of 56.47. In contrast, full SFT and LoRA often damage MATH-500 and IFEval performance. Reusing the memory across four Qwen3 backbones requires 9.23 EFLOPs, or 0.22 times the estimated full-SFT cost. The router is trained with domain data and general examples from NVIDIA’s Nemotron-Post-Training-Dataset-v1, making specialization modular without repeatedly retraining the backbone.

Original abstract

Adapting Large Language Models (LLMs) to specialized domains often incurs an alignment tax, as fine-tuning on domain-specific tasks can cause catastrophic forgetting and substantially degrade performance on general tasks. We propose MemSFT, which mitigates the alignment tax by decoupling domain specialization from backbone parameter updates through a plug-and-play parametric memory. The memory is trained to imitate the behavior of a non-parametric retriever operating over domain data, thereby memorizing knowledge and patterns that would otherwise be accessed through retrieval. Once trained on a specific domain, the memory can be reused across LLMs of different sizes. During generation, a learned router dynamically fuses the output distributions of the memory and backbone at each decoding step, allowing domain expertise to be invoked selectively. Across biology, geoscience, and law, evaluations with models ranging from Qwen3-8B to Qwen3-235B-A22B show that MemSFT consistently improves domain performance with negligible degradation in general performance, whereas full SFT suffers severe forgetting on general tasks. Overall, our results demonstrate a practical path to decoupling general model capabilities from domain-specific knowledge at the parameter level, thereby equipping LLMs with new specialized capabilities without compromising their general capabilities.

Read the original paper

More in Continual Learning

Browse all 24 papers →
02Continual Learning

From Knowledge Access to Source Learning: Developing Source-Specific Competence

Lucheng Fu, Kejing Xia, Yiyang Wang, Yiqiao Jin, Jinjin He, Xiyuan Yang, Haoxin Liu, Ye Yu, Haibo Jin, Yijia Xiao, Wenke Lee, B. Aditya Prakash, Haohan Wang

SourceLearn helps LLM agents progressively build reusable expertise about trusted information sources instead of repeatedly starting from scratch.

Read analysis
03Continual Learning

Local Support Learning

Assaf Ben-Kish, Akarsh Kumar, James Glass, Raja Giryes

Local Support Learning helps large language models learn new skills without overwriting what they already know by activating updates only where they are locally needed.

Read analysis