PEFT-Arena: Understanding Parameter-Efficient Finetuning from a Stability-Plasticity Perspective
AuthorsYangyi Huang, Ruotian Peng, Zeju Qiu, Jiale Kang, Yandong Wen, Bernhard Schölkopf, Weiyang Liu
Resources
PEFT-Arena reframes LLM finetuning as a balance between learning new tasks and preserving old abilities, revealing which parameter-efficient methods best avoid forgetting.
Key results
On Qwen2.5-7B under SFT, full finetuning raises the math target average from 35.30 to 50.63.
On Qwen2.5-7B under SFT, the general score drops from 46.97 to 34.22 after full finetuning.
Under SFT on Qwen2.5-7B, OFT-b32 reaches 46.93 on math with a much smaller general drop than full finetuning.
For Qwen2.5-7B SFT, OFT-b32 reduces the math general score by only 2.60 points relative to the base model.
For Qwen2.5-7B SFT with OFT-b32, Cayley-generator interpolation reaches 45.77 math at α = 0.3.
Across SFT checkpoints, Procrustes residual on general data correlates with forgetting at Pearson 0.711.
What the paper found
PEFT-Arena reframes parameter-efficient finetuning of large language models as a stability-plasticity problem, using Qwen2.5-7B and Llama3.2-3B-Instruct on math and medical reasoning while jointly scoring target accuracy and retained general ability on IFEval, Natural Questions, and BBH. Across supervised fine-tuning, full finetuning delivers the largest task gains but also the largest forgetting, with Qwen math rising from 35.30 to 50.63 while general performance falls from 46.97 to 34.22; by contrast, orthogonal finetuning, or OFT, sits on the strongest Pareto frontier under similar parameter budgets, such as OFT-b32 reaching 46.93 on math with only a 2.60-point general drop, and remaining competitive under GRPO-based RLVR. The paper explains this advantage geometrically: in weight space, OFT better preserves the pretrained singular-value structure, while methods like PiSSA and MiSS induce spiky updates and larger spectral disruption; in activation space, forgetting correlates with non-isometric distortion, with Procrustes residual correlating 0.711 with forgetting and PiSSA showing extreme distortion and 34.56 points of forgetting. Interpolation reveals that many SFT checkpoints overshoot the best trade-off, and parameterization-aware paths matter: scaling OFT’s Cayley generator outperforms naive dense-weight interpolation, improving Qwen math to 45.77 and general to 48.64 at α=0.3. The authors also show that layer-wise OFT rewinding can recover a better retention point, suggesting post-hoc geometry-aware control as a practical mitigation for forgetting.
Original abstract
Parameter-efficient finetuning (PEFT) has become the standard approach for adapting large language models, yet evaluations largely emphasize downstream accuracy while overlooking the retention of pretrained capabilities. We argue that PEFT should be assessed through the stability-plasticity dilemma: the trade-off between target-task adaptation and resistance to forgetting. We introduce PEFT-Arena, a benchmark that jointly measures downstream performance and general capability retention. Across methods, we find distinct stability-plasticity profiles; under comparable parameter budgets, orthogonal finetuning achieves the most favorable Pareto frontier. To explain these differences, we analyze PEFT updates from two geometric perspectives. In weight space, spectral analysis reveals how parameterizations interact with the pretrained singular-value structure. In activation space, retention metrics show whether finetuning preserves or distorts general-capability representations, with forgetting linked to non-isometric representation distortion. Finally, an analysis shows that final SFT checkpoints often overshoot a better target-retention operating point. Inspired by this, we present case studies of a post-hoc improvement with path-wise rewinding.
Read the original paperMore in Large Language Models
Browse all 81 papers →Finetuning with Sampling: SFT Learns Better Than You Think
Aayush Karan, Sitan Chen, Yilun Du
By sampling and reshaping expert data before training, this work argues that supervised finetuning can match RL while generalizing better and forgetting less.
Generalization Dynamics of LM Pre-training
Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen
Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.
Rethinking Self-Distillation for Multi-Teacher Capability Merging
Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra
The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.