NTH

PEFT-Arena: Understanding Parameter-Efficient Finetuning from a Stability-Plasticity Perspective

AuthorsYangyi Huang, Ruotian Peng, Zeju Qiu, Jiale Kang, Yandong Wen, Bernhard Schölkopf, Weiyang Liu

June 6, 2026 2 min read
Watch on YouTube
The one-line take

PEFT-Arena reframes LLM finetuning as a balance between learning new tasks and preserving old abilities, revealing which parameter-efficient methods best avoid forgetting.

Key results

50.63
Qwen math target gain (Full FT)

On Qwen2.5-7B under SFT, full finetuning raises the math target average from 35.30 to 50.63.

34.22
Qwen math general after Full FT

On Qwen2.5-7B under SFT, the general score drops from 46.97 to 34.22 after full finetuning.

46.93
Qwen OFT-b32 math target

Under SFT on Qwen2.5-7B, OFT-b32 reaches 46.93 on math with a much smaller general drop than full finetuning.

2.60
Qwen OFT-b32 general drop

For Qwen2.5-7B SFT, OFT-b32 reduces the math general score by only 2.60 points relative to the base model.

45.77
OFT-Cayley math at alpha 0.3

For Qwen2.5-7B SFT with OFT-b32, Cayley-generator interpolation reaches 45.77 math at α = 0.3.

0.711
Procrustes residual correlation

Across SFT checkpoints, Procrustes residual on general data correlates with forgetting at Pearson 0.711.

What the paper found

PEFT-Arena reframes parameter-efficient finetuning of large language models as a stability-plasticity problem, using Qwen2.5-7B and Llama3.2-3B-Instruct on math and medical reasoning while jointly scoring target accuracy and retained general ability on IFEval, Natural Questions, and BBH. Across supervised fine-tuning, full finetuning delivers the largest task gains but also the largest forgetting, with Qwen math rising from 35.30 to 50.63 while general performance falls from 46.97 to 34.22; by contrast, orthogonal finetuning, or OFT, sits on the strongest Pareto frontier under similar parameter budgets, such as OFT-b32 reaching 46.93 on math with only a 2.60-point general drop, and remaining competitive under GRPO-based RLVR. The paper explains this advantage geometrically: in weight space, OFT better preserves the pretrained singular-value structure, while methods like PiSSA and MiSS induce spiky updates and larger spectral disruption; in activation space, forgetting correlates with non-isometric distortion, with Procrustes residual correlating 0.711 with forgetting and PiSSA showing extreme distortion and 34.56 points of forgetting. Interpolation reveals that many SFT checkpoints overshoot the best trade-off, and parameterization-aware paths matter: scaling OFT’s Cayley generator outperforms naive dense-weight interpolation, improving Qwen math to 45.77 and general to 48.64 at α=0.3. The authors also show that layer-wise OFT rewinding can recover a better retention point, suggesting post-hoc geometry-aware control as a practical mitigation for forgetting.

Original abstract

Parameter-efficient finetuning (PEFT) has become the standard approach for adapting large language models, yet evaluations largely emphasize downstream accuracy while overlooking the retention of pretrained capabilities. We argue that PEFT should be assessed through the stability-plasticity dilemma: the trade-off between target-task adaptation and resistance to forgetting. We introduce PEFT-Arena, a benchmark that jointly measures downstream performance and general capability retention. Across methods, we find distinct stability-plasticity profiles; under comparable parameter budgets, orthogonal finetuning achieves the most favorable Pareto frontier. To explain these differences, we analyze PEFT updates from two geometric perspectives. In weight space, spectral analysis reveals how parameterizations interact with the pretrained singular-value structure. In activation space, retention metrics show whether finetuning preserves or distorts general-capability representations, with forgetting linked to non-isometric representation distortion. Finally, an analysis shows that final SFT checkpoints often overshoot a better target-retention operating point. Inspired by this, we present case studies of a post-hoc improvement with path-wise rewinding.

Read the original paper

More in Large Language Models

Browse all 81 papers →
02Llm

Generalization Dynamics of LM Pre-training

Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen

Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.

Read analysis
03Llm

Rethinking Self-Distillation for Multi-Teacher Capability Merging

Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra

The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.

Read analysis