Decoupling Exploration from Optimization in RLVR
- Published
- Source
- arXiv
- Paper number
- 1177
- Field
- Agents
- arXiv ID
- 2610.10536
Key points
- They confirmed that adding a novelty bonus directly to RLVR degrades general capabilities outside the supervised reward signal.
- They proposed ExpDis, which separates an explorer model from a student model and distills only the correct solutions filtered from the explorer.
- Across seven math benchmarks it beat DAPO in both average accuracy and pass@k under the same compute budget.
- It showed continued gains as parallel explorers and repeated rounds were scaled independently.
Paper links
External research summaries. These are not HDATF publications or measured product results.