Understanding Evolution Strategies for LLM Reasoning: Broader Reasoning Coverage than GRPO
- Published
- Source
- arXiv
- Paper number
- 1042
- Field
- Machine Learning
- arXiv ID
- 2608.27351
Key points
- It showed that, unlike GRPO, ES improves Pass@1 and Pass@K together without entropy collapse (loss of exploration diversity), preserving diversity in reasoning answers.
- A sequential combination that first trains with GRPO and then finishes with ES combined the strengths of both methods.
- Performance gains were concentrated in a small subset of large updates (mainly normalization and attention parameters), showing that functional changes are sparse even when parameters move substantially.
- ES largely preserved existing capabilities in held-out evaluations, providing a counterexample to the idea that large parameter movements themselves cause catastrophic forgetting.
- Z-score reward normalization is essential, and as model size increased, a population of just N=16 approached the performance of N=64.
Paper links
External research summaries. These are not HDATF publications or measured product results.