ExpRL: Exploratory RL for LLM Mid-Training
- Published
- Source
- arXiv
- Paper number
- 436
- Field
- Machine Learning
- arXiv ID
- 2606.17024
Key points
- It proposes two ExpRL variants, ExpRL-Outcome and ExpRL-Process, that use reference answers as scoring rubrics for an LLM judge rather than as imitation targets.
- On Qwen3-4B, ExpRL-Outcome reaches 34.2 percent pass@1 on AIME-25, far exceeding GRPO at 27.1 percent and SFT at 5.7 percent.
- The improvement in pass@k confirms that coverage, meaning the diversity of solution strategies, also improves.
- It observes more reasoning behaviors, such as verification, self-correction, and backtracking, than in the base model.
- For judges of 4B or larger, the matched reference for a problem shows the lowest misplacement rate, confirming that the judge validates based on the reference.
- Mixed-domain experiments across math, coding, and science suggest that ExpRL can scale beyond a single domain.
Paper links
External research summaries. These are not HDATF publications or measured product results.