Spectral Rewiring for Exploration, Purification, and Model Merging
AuthorsZhilong Zhang, Hongli Yu, Huan-ang Gao, Hanlin Wu, Yuxuan Song, Wei-Ying Ma, Ya-Qin Zhang, Hao Zhou
Resources
SAR trims language-model updates down to their most useful spectral directions, aiming to preserve reasoning while improving exploration, capability consolidation, and model merging.
Key results
Smallest SAR rewiring representation reported relative to total model parameters.
SAR preserves more than this fraction of post-training reasoning performance.
Coverage after 256 rollouts with SAR, versus 83.33% for full RL.
Average improvement across seven open agentic coding benchmarks.
Mix-RL performance after SAR projection, up from 67.61% for full RL.
AIME AVG@32 for the SAR math-code merged model.
What the paper found
Researchers from Tsinghua AIR and ByteDance Seed propose Subspace-Aligned Rewiring, or SAR, a training-free method for editing reinforcement-learning updates in large language models. SAR first extracts a low-rank component from the dense update, then projects it into the pretrained model’s singular-vector basis, producing a rewiring matrix whose off-diagonal terms reconnect latent capabilities for compositional reasoning while filtering off-manifold directions that cause exploration collapse and domain interference. Across DeepScaleR, POLARIS, and OLMo-3.1-32B-Think, spanning 1.5B to 32B parameters, SAR preserves more than 99% of peak reasoning performance with as little as 0.58% of total parameters. On AIME 2024 with 256 rollouts, it raises solution coverage to 86.66%, compared with 83.33% for full RL, showing improved high-k exploration without retraining. On an in-house coding model, a top-1% spectral projection improves six of seven agentic coding benchmarks, with a +2.52% average gain. In Mix-RL, SAR raises LiveCodeBench v5 from 67.61% to 69.09% while maintaining mathematics and instruction following. For math-code expert merging, it achieves 43.44% AIME AVG@32 and 32.25% LiveCodeBench AVG@8 at the 1.5B scale, exceeding the best individual experts simultaneously. The results position spectral alignment, rather than low-rank compression alone, as the key mechanism for purification, exploration, and cross-domain model merging.
Original abstract
Reinforcement learning has become a standard post-training recipe for large language models, but dense full-parameter updates create two deployment-relevant bottlenecks: suppressed reasoning performance, often reflected by premature saturation of test-time scaling, and interference when consolidating multiple capabilities through multi-domain training or model merging. We show that the reasoning-effective component of these updates is largely concentrated in the base model's spectral space, motivating Subspace-Aligned Rewiring (SAR), a post-hoc editing method that retains this spectral core while removing orthogonal components. SAR therefore preserves reasoning gains and filters residual update directions that suppress performance or amplify cross-domain interference. Across several model families and scales, SAR extracts compact reasoning cores using as little as approximately 0.58% of total parameters: it preserves over 99% of post-training performance and improves high-k exploration in mathematical reasoning, and generalizes to agentic coding by improving six of seven open benchmarks on an in-house model. SAR also purifies mixed-domain training updates by releasing suppressed coding capability while maintaining math reasoning and instruction following. It further enables model merging across experts, yielding cross-domain generalization that surpasses previous merging baselines and even the best single-domain experts. Overall, SAR shows that extracting reasoning-effective updates from parameter geometry can serve as a training-free mechanism to improve reasoning and multi-domain performance.
Read the original paperMore in Large Language Models
Browse all 81 papers →Finetuning with Sampling: SFT Learns Better Than You Think
Aayush Karan, Sitan Chen, Yilun Du
By sampling and reshaping expert data before training, this work argues that supervised finetuning can match RL while generalizing better and forgetting less.
Generalization Dynamics of LM Pre-training
Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen
Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.
Rethinking Self-Distillation for Multi-Teacher Capability Merging
Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra
The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.