Decoupling Exploration from Optimization in RLVR

Published
Source
arXiv
Paper number
1177
Field
Agents
arXiv ID
2610.10536

Key points

  • They confirmed that adding a novelty bonus directly to RLVR degrades general capabilities outside the supervised reward signal.
  • They proposed ExpDis, which separates an explorer model from a student model and distills only the correct solutions filtered from the explorer.
  • Across seven math benchmarks it beat DAPO in both average accuracy and pass@k under the same compute budget.
  • It showed continued gains as parallel explorers and repeated rounds were scaled independently.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)