Breaking Entropy Bounds: Accelerating RL Training via MTP with Rejection Sampling
- Published
- Source
- arXiv
- Paper number
- 396
- Field
- Machine Learning
- arXiv ID
- 2606.12370
Key points
- We found a clear negative linear relationship between MTP acceptance rate and policy entropy: as entropy rises, acceptance rate falls linearly.
- We mathematically and empirically show that probabilistic rejection sampling is far more robust to entropy variation than greedy target-only sampling.
- Instead of conventional CE/KL losses, we propose an end-to-end TV loss that directly optimizes rejection-sampling acceptance rate, improving it by about 10 percentage points.
- A light MTP pretraining phase before RL is sufficient to keep acceptance rates stable throughout RL, eliminating the need for online MTP co-training.
- Across mathematical reasoning, code generation, and agentic tasks, it reaches up to 95% acceptance and an additional 25% increase in reasoning throughput.
- In asynchronous RL pipelines for Qwen3.5/3.6/3.7, it achieves up to 1.8x end-to-end acceleration.
Paper links
External research summaries. These are not HDATF publications or measured product results.