ExpRL: Exploratory RL for LLM Mid-Training

Published
Source
arXiv
Paper number
436
Field
Machine Learning
arXiv ID
2606.17024

Key points

  • It proposes two ExpRL variants, ExpRL-Outcome and ExpRL-Process, that use reference answers as scoring rubrics for an LLM judge rather than as imitation targets.
  • On Qwen3-4B, ExpRL-Outcome reaches 34.2 percent pass@1 on AIME-25, far exceeding GRPO at 27.1 percent and SFT at 5.7 percent.
  • The improvement in pass@k confirms that coverage, meaning the diversity of solution strategies, also improves.
  • It observes more reasoning behaviors, such as verification, self-correction, and backtracking, than in the base model.
  • For judges of 4B or larger, the matched reference for a problem shows the lowest misplacement rate, confirming that the judge validates based on the reference.
  • Mixed-domain experiments across math, coding, and science suggest that ExpRL can scale beyond a single domain.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)