TREK: Distill to Explore, Reinforce to Refine
- Published
- Source
- arXiv
- Paper number
- 580
- Field
- Machine Learning
- arXiv ID
- 2607.05339
Key points
- It diagnoses a limitation of GRPO: the bottleneck is hard prompts for which the student never samples a reward trajectory.
- The output-only interface requires only validated trajectories, not the teacher's logits or internal states, so it is compatible with black-box teachers.
- Prompt routing queries the teacher only for prompts whose student pass rate pS(x) is less than or equal to tau_low, and it selects the top-r student-near trajectories.
- In math tasks, Qwen3-8B improves from 36.9 to 40.3 on AIME 2025 and from 47.9 to 51.1 on AIME 2024, both averaged at 16 samples, and the self-context variant also improves meaningfully.
- In agent tasks, it improves from 75.8 to 82.8 on ALFWorld and from 12.5 to 26.7 on ScienceWorld, which shows especially large gains on difficult tasks.
- It returns to standard GRPO after the forward-KL integration stage, where the validated mode is further refined and strengthened.
Paper links
External research summaries. These are not HDATF publications or measured product results.