Understanding Evolution Strategies for LLM Reasoning: Broader Reasoning Coverage than GRPO
AuthorsYunpeng Ba, Zhi Zheng, Yue Xie, Jiaqing Li, Xialiang Tong, Tao Zhong, Mingxuan Yuan, Zhichao Lu, Xuyang Wu, Zhenkun Wang
Resources
This work argues that Evolution Strategies can preserve broader reasoning diversity than GRPO, improving LLM Pass@K performance while complementing GRPO's strength in Pass@1.
Key results
DeepSeek-R1-Distill-Qwen-1.5B trained on DeepScaleR
Easy-Setting comparisons where GRPO falls below the base model on both Pass@16 and Pass@32
Qwen2.5-7B-Instruct relative whole-model L2 distance
Llama-3.2-3B-Instruct changed coordinates removable while retaining most ES performance
Within 0.01 of the N=64 reward reference for Qwen2.5-1.5B-Instruct and Qwen2.5-3B-Instruct
What the paper found
This study compares memory-efficient Evolution Strategies, or ES, with Group Relative Policy Optimization, GRPO, for verifier-guided reasoning in LLMs. ES perturbs model parameters, evaluates the resulting population with forward-only rollouts, and aggregates z-score-normalized rewards, while GRPO backpropagates token-level advantages through one policy. Across OpenAI’s GSM8K and DeepScaleR, ES improves Pass@1 while preserving broader reasoning coverage at larger sampling budgets: on DeepSeek-R1-Distill-Qwen-1.5B, ES reaches average Pass@1, Pass@16, and Pass@32 scores of 49.9, 75.0, and 78.9 percent, compared with GRPO’s 52.9, 74.7, and 78.0 percent. GRPO falls below its base model on Pass@16 and Pass@32 in 15 of 18 Easy-Setting comparisons, consistent with entropy collapse, whereas ES maintains more stable entropy. The analysis attributes ES’s Pass@K advantage to verifier-projected Jensen–Shannon diversity among perturbed policies: heterogeneous population members provide wider access to correct reasoning paths. Sequential ES→GRPO and GRPO→ES training adds useful Pass@1–Pass@K trade-offs. Although ES moves models as far as 44.1 times farther from initialization than GRPO, performance gains are functionally sparse; removing updates covering up to 92.64 percent of changed coordinates can largely preserve accuracy, with important updates concentrated in LayerNorm and attention parameters. Finally, z-score normalization is essential, two-point estimation offers no advantage for regenerated reasoning rollouts, and larger Qwen2.5 models can use smaller populations: N=16 is within 0.01 of the N=64 reward reference for both 1.5B and 3B models. The results position ES as a distinct reasoning post-training paradigm rather than merely a cheaper GRPO substitute.
Original abstract
Evolution Strategies (ES) have recently emerged as a memory-efficient post-training paradigm for LLM reasoning. However, the optimization behavior of ES remains understudied, making it hard to define its advantage scope compared to mainstream post-training paradigms (e.g., Group Relative Policy Optimization (GRPO)). By systematically investigating ES dynamics and mechanisms, this paper first identifies a performance advantage of ES over GRPO, theoretically and empirically showing that ES can lead to broader reasoning coverage, thereby better exploiting the reasoning capabilities of pretrained LLMs. Theoretically, we show that verifier-projected Jensen-Shannon diversity across the ES population is helpful to higher Pass@K performances. Empirically, unlike GRPO, which exhibits entropy collapse, ES improves Pass@1 while attaining higher Pass@K than GRPO. We further develop a sequential GRPO-ES training strategy that combines GRPO's strength in Pass@1 with ES's gains in Pass@K. Second, we find that despite substantial whole-model parameter drift, the task-performance gains of ES are only contributed to a sparse subset of larger-magnitude updates. This functional sparsity suggests that large parameter movement need not imply widespread functional change, and held-out evaluations further show that it does not necessarily lead to catastrophic forgetting. Finally, we study how hyperparameter design affects the effectiveness of ES, demonstrating that ES requires a smaller population size in a larger LLM. These findings position ES as a distinct reasoning post-training paradigm rather than a less effective, memory-efficient alternative to GRPO.
Read the original paperMore in AI Reasoning
Browse all 39 papers →Knowing When Thinking Is Not Enough: Teaching Small Reasoning Models to Reason Beyond Their Parametric Knowledge
Chanuk Lee, Minki Kang, Sangwoo Park, Woongyeong Yeo, Jinheon Baek, Sung Ju Hwang
FlyBy teaches small reasoning models to recognize when more internal thinking will not help and instead ask a stronger model for missing knowledge.
On Language Drift during RLVR Post-Training
Michael Sullivan, Alexander Koller
RLVR can make reasoning models increasingly use strange internal languages, and preventing that drift may require sacrificing some performance.
Principled Thoughts for Latent Recursive LLM Systems
Fahd Seddik, Fatemeh Fard
REST teaches latent LLM agents to form more causal, minimal, separable, and stable internal thoughts, improving reasoning accuracy and interpretability.