NTH

Predictable GRPO: A Closed-Form Model of Training Dynamics

AuthorsRajat Ghosh, Datta Nimmaturi, Aryan Singhal, Vaishnavi Bhargava, Henry Wong, Johnu George, Debojyoti Dutta

July 1, 2026 3 min read
Watch on YouTube
The one-line take

This paper turns GRPO training into a predictable dynamical system, giving a closed-form explanation for reward growth, stability, and failure modes in LLM reasoning training.

Key results

0.91
training-fit R2

closed-form three-phase law fit to GSM8K training reward across three models and two group sizes

4
group sizes

one of the two GRPO rollout group sizes tested, compared against 16

16
group sizes

one of the two GRPO rollout group sizes tested, compared against 4

8
exact softmax completions

finite completion set used to validate the oscillatory regime and stability threshold exactly

0.12
measured stiffness

independently measured curvature scale used to predict the refresh-interval threshold

What the paper found

Predictable GRPO, from researchers at Nutanix and Unsloth, turns Group Relative Policy Optimization from an empirical curve-fitting exercise into a closed-form dynamical model. Under a single mean-field assumption, the paper reduces GRPO to a stochastically forced damped oscillator whose mass, damping, and stiffness are determined by optimizer hyperparameters and one measured curvature scale: momentum supplies inertia, stale-policy refresh erodes damping, and group size G acts only as noise temperature. This reframes the widely used saturating reward law as the overdamped limit of the second-order system, while the retained inertial term explains the slow-start phase that a single exponential cannot capture. On GSM8K, three open-weight models—Nemotron-Mini-4B, DeepSeek-LLM-7B-Chat, and DeepSeek-R1-Distill-Qwen-1.5B—trained with GRPO at G = 4 and G = 16 are fit by the three-phase law with R2 ≥ 0.91, and the deterministic trajectory is approximately G-invariant while stationary fluctuations shrink as 1/G. The model also predicts a refresh-interval stability threshold, KηKref > 1 − µ, and a monotone-to-oscillatory transition; these are confirmed in an exact softmax-bandit reduction with p = Eπθ[R] closed exactly, using a single prompt with 8 completions and a measured stiffness of 0.12. Out of distribution, GSM8K-trained policies transfer to eight math benchmarks, including GSM-Plus, MetaMathQA, OpenMath2, NuminaMath, MATH-500, IMO-Bench, AIME-2026, and HMMT-2025, with the strongest gains near the training distribution and much weaker gains on competition-level tasks.

Original abstract

Group Relative Policy Optimization (GRPO) has become a standard tool for improving the reasoning ability of large language models, yet its training dynamics are still described empirically: reward trajectories are fit with low-parameter functional forms whose constants carry no mechanistic meaning, and hyperparameter choices remain a matter of trial and error. We develop a first-principles reduced-order model of these dynamics. The reduction has three consequences. First, it subsumes the empirical single-exponential saturation law as its overdamped limit, recasting the fitted plateau, timescale, and size exponent as the fixed point, inverse stiffness, and curvature-scaling exponent of the underlying potential, and adding, through the retained inertial term, the slow-start phase the single exponential cannot represent. Second, it yields predictions tied to independently measurable quantities rather than fitted ones: group-size invariance of the deterministic trajectory with a $1/G$ stationary fluctuation, a sharp stability threshold in the refresh interval, and an overdamped-to-oscillatory transition. Third, it furnishes diagnostics that separate failure modes a reward curve alone conflates -- reward hacking, advantage degeneracy, policy concentration, and dynamical instability. Across three models and two group sizes, the closed-form trajectory fits training reward to $R^2 \geq 0.91$ and the predicted group-size invariance holds on both the reward curve and out-of-distribution transfer to eight math benchmarks. The stability and oscillatory predictions are exercised in a controlled exact-reduction setting where the mean-field assumption holds exactly: a softmax-bandit reduction reproduces the predicted overdamped-to-oscillatory transition and locates the refresh-interval stability threshold at the independently measured stiffness, with a deep-network demonstration left to future work.

Read the original paper

More in Large Language Models

Browse all 81 papers →
02Llm

Generalization Dynamics of LM Pre-training

Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen

Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.

Read analysis
03Llm

Rethinking Self-Distillation for Multi-Teacher Capability Merging

Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra

The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.

Read analysis