NTH

Understanding Evolution Strategies for LLM Reasoning: Broader Reasoning Coverage than GRPO

AuthorsYunpeng Ba, Zhi Zheng, Yue Xie, Jiaqing Li, Xialiang Tong, Tao Zhong, Mingxuan Yuan, Zhichao Lu, Xuyang Wu, Zhenkun Wang

September 6, 2026 3 min read
Watch on YouTube
The one-line take

This work argues that Evolution Strategies can preserve broader reasoning diversity than GRPO, improving LLM Pass@K performance while complementing GRPO's strength in Pass@1.

Key results

78.9
Hard-setting ES average Pass@32

DeepSeek-R1-Distill-Qwen-1.5B trained on DeepScaleR

15
GRPO large-K degradation frequency

Easy-Setting comparisons where GRPO falls below the base model on both Pass@16 and Pass@32

44.1
Maximum ES-to-GRPO parameter drift ratio

Qwen2.5-7B-Instruct relative whole-model L2 distance

92.64%
Update sparsity threshold

Llama-3.2-3B-Instruct changed coordinates removable while retaining most ES performance

N=16
Reduced population size

Within 0.01 of the N=64 reward reference for Qwen2.5-1.5B-Instruct and Qwen2.5-3B-Instruct

What the paper found

This study compares memory-efficient Evolution Strategies, or ES, with Group Relative Policy Optimization, GRPO, for verifier-guided reasoning in LLMs. ES perturbs model parameters, evaluates the resulting population with forward-only rollouts, and aggregates z-score-normalized rewards, while GRPO backpropagates token-level advantages through one policy. Across OpenAI’s GSM8K and DeepScaleR, ES improves Pass@1 while preserving broader reasoning coverage at larger sampling budgets: on DeepSeek-R1-Distill-Qwen-1.5B, ES reaches average Pass@1, Pass@16, and Pass@32 scores of 49.9, 75.0, and 78.9 percent, compared with GRPO’s 52.9, 74.7, and 78.0 percent. GRPO falls below its base model on Pass@16 and Pass@32 in 15 of 18 Easy-Setting comparisons, consistent with entropy collapse, whereas ES maintains more stable entropy. The analysis attributes ES’s Pass@K advantage to verifier-projected Jensen–Shannon diversity among perturbed policies: heterogeneous population members provide wider access to correct reasoning paths. Sequential ES→GRPO and GRPO→ES training adds useful Pass@1–Pass@K trade-offs. Although ES moves models as far as 44.1 times farther from initialization than GRPO, performance gains are functionally sparse; removing updates covering up to 92.64 percent of changed coordinates can largely preserve accuracy, with important updates concentrated in LayerNorm and attention parameters. Finally, z-score normalization is essential, two-point estimation offers no advantage for regenerated reasoning rollouts, and larger Qwen2.5 models can use smaller populations: N=16 is within 0.01 of the N=64 reward reference for both 1.5B and 3B models. The results position ES as a distinct reasoning post-training paradigm rather than merely a cheaper GRPO substitute.

Original abstract

Evolution Strategies (ES) have recently emerged as a memory-efficient post-training paradigm for LLM reasoning. However, the optimization behavior of ES remains understudied, making it hard to define its advantage scope compared to mainstream post-training paradigms (e.g., Group Relative Policy Optimization (GRPO)). By systematically investigating ES dynamics and mechanisms, this paper first identifies a performance advantage of ES over GRPO, theoretically and empirically showing that ES can lead to broader reasoning coverage, thereby better exploiting the reasoning capabilities of pretrained LLMs. Theoretically, we show that verifier-projected Jensen-Shannon diversity across the ES population is helpful to higher Pass@K performances. Empirically, unlike GRPO, which exhibits entropy collapse, ES improves Pass@1 while attaining higher Pass@K than GRPO. We further develop a sequential GRPO-ES training strategy that combines GRPO's strength in Pass@1 with ES's gains in Pass@K. Second, we find that despite substantial whole-model parameter drift, the task-performance gains of ES are only contributed to a sparse subset of larger-magnitude updates. This functional sparsity suggests that large parameter movement need not imply widespread functional change, and held-out evaluations further show that it does not necessarily lead to catastrophic forgetting. Finally, we study how hyperparameter design affects the effectiveness of ES, demonstrating that ES requires a smaller population size in a larger LLM. These findings position ES as a distinct reasoning post-training paradigm rather than a less effective, memory-efficient alternative to GRPO.

Read the original paper

More in AI Reasoning

Browse all 39 papers →
02Reasoning

On Language Drift during RLVR Post-Training

Michael Sullivan, Alexander Koller

RLVR can make reasoning models increasingly use strange internal languages, and preventing that drift may require sacrificing some performance.

Read analysis