A Local Perturbation Theory for Cross-Domain Interference and Recovery in Multi-Domain RL
AuthorsLei Yang, Siyu Ding, Deyi Xiong
Resources
This paper explains why training an LLM on one RL domain can hurt others, and shows that a small targeted refresh can recover lost performance with minimal side effects.
Key results
Math score after sequential Code → Math → QA → CW training
Math score immediately after Math training in the sequential curriculum
Approximate share of parameters with negligible absolute change in single-domain RL
Upper end of the reported share of parameters with negligible absolute change
Math score after short refresh from CWo
Best four-domain average score after refresh
What the paper found
This paper from Tianjin University and Baidu Inc. argues that cross-domain interference in multi-domain RL for LLM post-training is a localized route-level phenomenon rather than a global gradient-conflict problem. Using Qwen3-4B-Thinking-2507 and GRPO in VeRL, the authors train four domains—Math, Code, QA, and creative writing—and show that sequential Code → Math → QA → CW training can drop Math from 66.49 to 57.66 even while whole-model gradient cosines stay near zero. Their structural analysis finds that single-domain RL makes sparse, small-magnitude edits, with about 77%–89% of parameters changing by less than 1e-7, while top-changed neurons overlap weakly across domains with Jaccard coefficients below 0.19; however, reasoning domains still share active computation routes, and directional alignment on those shared routes determines whether updates are synergistic or conflicting. The theory formalizes this as second-order local damage concentrated in a low-dimensional shared conflict subspace, and proves that a short refresh on the damaged domain geometrically contracts the harmful component. Empirically, a brief Re-Math refresh after CWo restores Math to 66.04 and lifts the average score to 66.39, outperforming JT at 65.62 and CGPO at 65.30. A training-free rollback on a 2% proxy set of conflict coordinates recovers 20.4% of the QA-induced Math loss, and a larger joint MLP+attention rollback reaches 73.6% recovery at a 32% budget.
Original abstract
Reinforcement learning (RL) post-training improves large language models (LLMs) on individual domains such as mathematical reasoning, code generation, question answering, and creative writing (CW), but training on one domain often degrades performance on others. Existing explanations based on catastrophic forgetting or global gradient conflict are incomplete: substantial interference can occur even when full-model gradients are nearly orthogonal. We show that single-domain RL produces sparse, small-magnitude parameter edits with weak overlap among top-changed neurons, while different domains still share substantial active computation routes on which update directions determine whether they act synergistically or conflict. Guided by this observation, we prove under a local perturbation model of multi-domain RL that later-domain training harms an earlier domain mainly through a second-order damage term, which under the observed sparse route structure concentrates in a low-dimensional shared conflict subspace. Moreover, a short domain refresh contracts the harmful component on this subspace, enabling selective recovery with limited collateral damage. Consistent with the theory, a brief Re-Math refresh after Code $\rightarrow$ Math $\rightarrow$ QA $\rightarrow$ CW recovers Math from 57.66 to 66.04 while largely preserving performance on the other domains, yielding the best average score of 66.39. Beyond refresh, a training-free rollback on a sparse proxy conflict coordinate set for the Math-QA pair partially restores Math, providing direct proxy-level evidence for localized damage. These results provide a localized mechanistic account of interference and recovery in multi-domain RL.
Read the original paperMore in Reinforcement Learning
Browse all 54 papers →Res-HIL: Human-Guided Residual Reinforcement Learning for Sample-Efficient Dexterous Manipulation
Mariia Iavorskaia, Christian Dietz, Sebastian Albrecht, Majid Khadiv
Res-HIL lets humans efficiently improve robot manipulation skills by teaching a small corrective policy on top of an existing imitation policy.
Selecting Diverse SFT Traces Improves Post-RL Generalization
Dylan Zhang, Mingyuan Wu, Jinning Li
Choosing varied reasoning paths—not just correct ones—can make reinforcement-trained language models generalize better.
Verifiable Hidden Dynamics Play: Generating Agentic RL Environments from Solved Mechanisms
Xinjie Shen, Wei Fan, Xudong Guo, Jianhong Tu, Yang Su, Chuqiao Kuang, Yinger Zhang, Dayiheng Liu
VHD-Play turns solved mathematical mechanisms into cheap, stateful, self-verifying worlds where language-model agents can practice long-horizon decision-making.