Understanding Evolution Strategies for LLM Reasoning: Broader Reasoning Coverage than GRPO

Published
Source
arXiv
Paper number
1042
Field
Machine Learning
arXiv ID
2608.27351

Key points

  • It showed that, unlike GRPO, ES improves Pass@1 and Pass@K together without entropy collapse (loss of exploration diversity), preserving diversity in reasoning answers.
  • A sequential combination that first trains with GRPO and then finishes with ES combined the strengths of both methods.
  • Performance gains were concentrated in a small subset of large updates (mainly normalization and attention parameters), showing that functional changes are sparse even when parameters move substantially.
  • ES largely preserved existing capabilities in held-out evaluations, providing a counterexample to the idea that large parameter movements themselves cause catastrophic forgetting.
  • Z-score reward normalization is essential, and as model size increased, a population of just N=16 approached the performance of N=64.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)