TREK: Distill to Explore, Reinforce to Refine

Published
Source
arXiv
Paper number
580
Field
Machine Learning
arXiv ID
2607.05339

Key points

  • It diagnoses a limitation of GRPO: the bottleneck is hard prompts for which the student never samples a reward trajectory.
  • The output-only interface requires only validated trajectories, not the teacher's logits or internal states, so it is compatible with black-box teachers.
  • Prompt routing queries the teacher only for prompts whose student pass rate pS(x) is less than or equal to tau_low, and it selects the top-r student-near trajectories.
  • In math tasks, Qwen3-8B improves from 36.9 to 40.3 on AIME 2025 and from 47.9 to 51.1 on AIME 2024, both averaged at 16 samples, and the self-context variant also improves meaningfully.
  • In agent tasks, it improves from 75.8 to 82.8 on ALFWorld and from 12.5 to 26.7 on ScienceWorld, which shows especially large gains on difficult tasks.
  • It returns to standard GRPO after the forward-KL integration stage, where the validated mode is further refined and strengthened.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)