NTH

ISO: An RLVR-Native Optimization Stack

AuthorsHanqing Zhu, Wenyan Cong, Zhizhou Sha, Sagnik Mukherjee, Xinyuan Song, David González-Martínez, Xiaoxia Wu, Yuandong Tian, Shiwei Liu, David Z. Pan, Zhangyang "Atlas" Wang

July 25, 2026 3 min read
Watch on YouTube
The one-line take

ISO speeds up RLVR training by preserving a model’s singular-value spectrum while optimizing the input and output directions that encode new reasoning behaviors.

Key results

3%
RLVR spectral residual

Average fraction of checkpoint displacement unexplained by restoring the base spectrum.

1.8%
Fixed-spectrum reconstruction residual

Median unexplained update when retaining the incoming spectrum while adapting both singular frames.

0.509
Qwen3-8B ISO-AdamW accuracy

Aggregate accuracy reached after 210 training steps.

0.495
Qwen3-8B AdamW matched accuracy

AdamW reaches this aggregate accuracy after 270 steps, while ISO-AdamW reaches it after 100 steps.

63.80
Qwen2.5-7B ISO-Merger aggregate

Aggregate data-free merging score, compared with 62.88 for the strongest baseline.

What the paper found

The paper from researchers at The University of Texas at Austin, UIUC, Emory, Together AI, Recursive Superintelligence, and the ELLIS Institute Tübingen proposes ISO, an RLVR-native optimization stack based on spectral inheritance. Analyzing models including DeepSeek-R1-Distill-Qwen-1.5B and Qwen3, the authors find that reinforcement learning with verifiable rewards largely preserves a base model’s singular-value spectrum while changing both its left and right singular frames: on a long-horizon RLVR run, the spectral residual averages approximately 3% of total checkpoint displacement, and a fixed-spectrum reconstruction with both frames adaptable leaves only a 1.8% median residual. ISO-Optimizer parameterizes each weight as UΣ₀Vᵀ, freezes the base spectrum Σ₀, and applies optimizers such as AdamW or Muon to Stiefel-constrained frame variables with polar retraction. Across reasoning and coding experiments from 1.5B to 8B parameters on DeepMath-103K, ISO-AdamW reaches 0.495 on Qwen3-8B-Base in 100 steps, matching AdamW’s 0.495 at 270 steps, then reaches 0.509 at 210 steps. Its offline counterpart, ISO-Merger, combines shared-base specialists without new data, rollouts, gradients, or distillation; on Qwen2.5-7B-Instruct, it obtains an aggregate 63.80 versus 62.88 for the strongest data-free baseline. The central design principle is therefore to inherit the spectrum and optimize both singular frames, rather than reuse pre-training optimization directly for RLVR.

Original abstract

Reinforcement learning with verifiable rewards (RLVR) is rapidly advancing the reasoning capabilities of language models, yet the optimization layer that converts reward feedback into weight-space updates remains poorly understood. Building on our prior analysis (Zhu et al., 2025), we study this missing layer through the singular structure of model weights and identify spectral inheritance: RLVR can reuse the base model's weight spectra while acquiring new behavior through changes in the associated input and output singular frames. We operationalize spectral inheritance as Isospectral Optimization (ISO), an RLVR-native, fixed-spectrum optimization framework with complementary offline and online instantiations. Offline, ISO-Merger combines the frame changes of shared-base specialists into a single fixed-spectrum model, requiring no post-merge data, rollouts, gradient updates, or on-policy distillation (OPD). It recovers complementary specialist capabilities and achieves the strongest aggregate performance among the compared data-free merging methods. Online, ISO-Optimizer applies a chosen base optimizer, including AdamW and Muon, to the frame variables while keeping the base spectra fixed. Across reasoning and coding tasks ranging from 1.5B to 8B parameters, ISO-Optimizer improves accuracy in the reported runs and reaches matched scores with substantially fewer training steps. On Qwen3-8B-Base, AdamW reaches an aggregate accuracy of 0.495 after 270 training steps. ISO-AdamW reaches the same accuracy after only 100 training steps and improves further to 0.509 after 210 training steps. Together, ISO offers a concrete answer to RLVR's missing optimization layer: rather than inheriting pre-training optimization wholesale, design post-training around the structure of reward-driven adaptation: inherit the spectrum, optimize the frames.

Read the original paper

More in Optimization

Browse all 36 papers →