Omega-S: A Functional Resilience Index for LLM Fine-Tuning
AuthorsAlberto Acedo
Resources
Omega-S is a cheap weight-only regularizer that helps LLMs remember old skills while learning new ones, though it still needs broader testing.
Key results
Relative increase from 0.173 to 0.238 after sequential fine-tuning on Meta’s Llama-3-8B.
Retention after CodeAlpaca-20k to Wikitext-2 fine-tuning, compared with 62.9% without regularization.
Median reduction during training; this was the operative mechanism rather than clustering.
Latency overhead when Omega-S is applied every 10 steps.
The alignment channel led the magnitude channel on 8 of 10 seeds.
What the paper found
Omega-S is a drop-in regularizer for reducing catastrophic forgetting during LLM fine-tuning. Instead of storing previous-task weights or data like Elastic Weight Consolidation, it operates only on current weight matrices, constructing a graph from W Wᵀ and estimating Tr(A³) with Hutchinson probes at O(N²) cost. In a sequential CodeAlpaca-20k to Wikitext-2 experiment using Meta’s Llama-3-8B with rank-8 LoRA, retention was evaluated on 164 HumanEval problems across 10 seeds: absolute HumanEval pass@1 increased from 0.173 without regularization to 0.238 with Omega-S, a 37.7% relative gain, while the retention ratio rose from 62.9% to 84.1%. Omega-S outperformed tuned weight decay on 10 of 10 seeds and EWC on 8 of 10, although the baseline tuning protocol and high GPU nondeterminism limit the strength of those comparisons. The paper’s most important qualification is mechanistic: despite its topological objective, the logistic graph construction saturates clustering, so the observed effect comes almost entirely from reducing node-degree variance, which fell 5.87% during training; an ablation found the alignment channel stronger than the magnitude channel on 8 of 10 seeds. A contrast-preserving cosine formulation activated clustering but reduced retention on all 10 seeds. The penalty adds 3.7% single-GPU step latency when applied every 10 steps and requires no previous-task data, Fisher matrix, or archived weights.
Original abstract
Fine-tuning a large language model on new data degrades what it previously learned. We present Omega-S, a drop-in penalty computed from the weight matrix alone: it needs no previous-task data, no Fisher matrix and no stored copy of the old weights. It is three lines in an existing training loop and adds under 4% to the cost of a step. Retention. On Llama-3-8B with LoRA, fine-tuned from code to prose and measured by HumanEval over ten seeds, Omega-S retains more of the original capability than no regularisation on 9 of 10 seeds (0.173 -> 0.238 absolute pass@1; sign test one-sided p=0.011, Wilcoxon p=0.006), as a retention ratio, 62.9% -> 84.1%. It also beats tuned weight decay on 10 of 10 seeds (p=0.002) and tuned EWC on 8 of 10 (p=0.014), every arm re-measured in the same session. Mechanism, measured rather than asserted. Omega-S is topological by construction, its objective built from Tr(A^3), but we measured which of its four factors actually moves and three do not: their elasticity with respect to the weights is at or below 1e-4, against 9e-3 for the degree-variance term. As implemented, the composite reduces to a penalty on the variance of node degrees, which means row magnitude in square modules and directional alignment in non-square ones. We report this because a method whose name promises one thing and whose gradient does another should say so. We also enumerate the open design choices, including a contrast-preserving construction that does what it was designed to do and makes retention worse on all ten seeds. Repeating an identical configuration, same seed and same hardware, gives a standard deviation of 0.104 in retention ratio. We have not found this quantified for low-rank fine-tuning of language models, and it bounds every seed-paired comparison in this literature, ours included. Code, per-seed results and the full record of negative results are available.
Read the original paperMore in Continual Learning
Browse all 24 papers →ASCENT: Online Test-Time Training of Long-Horizon Agents via Self-Distillation of Verified Experience
Haodong Lu, Dong Gong
ASCENT lets deployed LLM agents learn from verified successes on the fly by converting hindsight about their own trajectories into lasting weight updates.
From Knowledge Access to Source Learning: Developing Source-Specific Competence
Lucheng Fu, Kejing Xia, Yiyang Wang, Yiqiao Jin, Jinjin He, Xiyuan Yang, Haoxin Liu, Ye Yu, Haibo Jin, Yijia Xiao, Wenke Lee, B. Aditya Prakash, Haohan Wang
SourceLearn helps LLM agents progressively build reusable expertise about trusted information sources instead of repeatedly starting from scratch.
Local Support Learning
Assaf Ben-Kish, Akarsh Kumar, James Glass, Raja Giryes
Local Support Learning helps large language models learn new skills without overwriting what they already know by activating updates only where they are locally needed.