Breaking Entropy Bounds: Accelerating RL Training via MTP with Rejection Sampling

Published
Source
arXiv
Paper number
396
Field
Machine Learning
arXiv ID
2606.12370

Key points

  • We found a clear negative linear relationship between MTP acceptance rate and policy entropy: as entropy rises, acceptance rate falls linearly.
  • We mathematically and empirically show that probabilistic rejection sampling is far more robust to entropy variation than greedy target-only sampling.
  • Instead of conventional CE/KL losses, we propose an end-to-end TV loss that directly optimizes rejection-sampling acceptance rate, improving it by about 10 percentage points.
  • A light MTP pretraining phase before RL is sufficient to keep acceptance rates stable throughout RL, eliminating the need for online MTP co-training.
  • Across mathematical reasoning, code generation, and agentic tasks, it reaches up to 95% acceptance and an additional 25% increase in reasoning throughput.
  • In asynchronous RL pipelines for Qwen3.5/3.6/3.7, it achieves up to 1.8x end-to-end acceleration.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)