ISO: An RLVR-Native Optimization Stack
AuthorsHanqing Zhu, Wenyan Cong, Zhizhou Sha, Sagnik Mukherjee, Xinyuan Song, David González-Martínez, Xiaoxia Wu, Yuandong Tian, Shiwei Liu, David Z. Pan, Zhangyang "Atlas" Wang
Resources
ISO speeds up RLVR training by preserving a model’s singular-value spectrum while optimizing the input and output directions that encode new reasoning behaviors.
Key results
Average fraction of checkpoint displacement unexplained by restoring the base spectrum.
Median unexplained update when retaining the incoming spectrum while adapting both singular frames.
Aggregate accuracy reached after 210 training steps.
AdamW reaches this aggregate accuracy after 270 steps, while ISO-AdamW reaches it after 100 steps.
Aggregate data-free merging score, compared with 62.88 for the strongest baseline.
What the paper found
The paper from researchers at The University of Texas at Austin, UIUC, Emory, Together AI, Recursive Superintelligence, and the ELLIS Institute Tübingen proposes ISO, an RLVR-native optimization stack based on spectral inheritance. Analyzing models including DeepSeek-R1-Distill-Qwen-1.5B and Qwen3, the authors find that reinforcement learning with verifiable rewards largely preserves a base model’s singular-value spectrum while changing both its left and right singular frames: on a long-horizon RLVR run, the spectral residual averages approximately 3% of total checkpoint displacement, and a fixed-spectrum reconstruction with both frames adaptable leaves only a 1.8% median residual. ISO-Optimizer parameterizes each weight as UΣ₀Vᵀ, freezes the base spectrum Σ₀, and applies optimizers such as AdamW or Muon to Stiefel-constrained frame variables with polar retraction. Across reasoning and coding experiments from 1.5B to 8B parameters on DeepMath-103K, ISO-AdamW reaches 0.495 on Qwen3-8B-Base in 100 steps, matching AdamW’s 0.495 at 270 steps, then reaches 0.509 at 210 steps. Its offline counterpart, ISO-Merger, combines shared-base specialists without new data, rollouts, gradients, or distillation; on Qwen2.5-7B-Instruct, it obtains an aggregate 63.80 versus 62.88 for the strongest data-free baseline. The central design principle is therefore to inherit the spectrum and optimize both singular frames, rather than reuse pre-training optimization directly for RLVR.
Original abstract
Reinforcement learning with verifiable rewards (RLVR) is rapidly advancing the reasoning capabilities of language models, yet the optimization layer that converts reward feedback into weight-space updates remains poorly understood. Building on our prior analysis (Zhu et al., 2025), we study this missing layer through the singular structure of model weights and identify spectral inheritance: RLVR can reuse the base model's weight spectra while acquiring new behavior through changes in the associated input and output singular frames. We operationalize spectral inheritance as Isospectral Optimization (ISO), an RLVR-native, fixed-spectrum optimization framework with complementary offline and online instantiations. Offline, ISO-Merger combines the frame changes of shared-base specialists into a single fixed-spectrum model, requiring no post-merge data, rollouts, gradient updates, or on-policy distillation (OPD). It recovers complementary specialist capabilities and achieves the strongest aggregate performance among the compared data-free merging methods. Online, ISO-Optimizer applies a chosen base optimizer, including AdamW and Muon, to the frame variables while keeping the base spectra fixed. Across reasoning and coding tasks ranging from 1.5B to 8B parameters, ISO-Optimizer improves accuracy in the reported runs and reaches matched scores with substantially fewer training steps. On Qwen3-8B-Base, AdamW reaches an aggregate accuracy of 0.495 after 270 training steps. ISO-AdamW reaches the same accuracy after only 100 training steps and improves further to 0.509 after 210 training steps. Together, ISO offers a concrete answer to RLVR's missing optimization layer: rather than inheriting pre-training optimization wholesale, design post-training around the structure of reward-driven adaptation: inherit the spectrum, optimize the frames.
Read the original paperMore in Optimization
Browse all 36 papers →An $Ω(κ_y^8ε^{-6})$ Lower Bound for Stochastic NC-SC Bilevel Optimization with First-order Oracles
Zhihao Gu, Qilong Wu, Junchi Yang
This work proves that stochastic bilevel optimization fundamentally requires up to epsilon^{-6} oracle queries, showing existing methods are asymptotically optimal.
Hyper Algorithm Design Agent: Evolving Learnable Optimizer from Zero
Zipei Yu, Yue-Jiao Gong, Zeyuan Ma, Yuncheng Jiang, Zhiguang Cao
A pair of self-improving coding agents evolves new learnable optimization algorithms from a simple template, reducing the need for handcrafted optimizer design.
Tight Regret Bound for Online Inverse Linear Optimization via Multiscale Matrix Weights
Shinsaku Sakaue
A new multiscale matrix-weights algorithm learns hidden linear preferences online with provably optimal dimension-dependent regret.