Reinforcement Learning from Rich Feedback with Distributional DAgger
- Published
- Source
- arXiv
- Paper number
- 311
- Field
- Machine Learning
- arXiv ID
- 2606.05152
Key points
- The paper shows that RL methods using existing self-distillation objectives based on reverse KL or Jensen-Shannon do not guarantee monotonic policy improvement, meaning the update can increase the probability of worse actions even when the expert has higher reward.
- In contrast, it shows that forward cross-entropy allows monotonic policy improvement and provides regret guarantees.
- It also shows that this objective improves Pass@N by optimizing a lower bound on the teacher's weighted success likelihood.
Paper links
External research summaries. These are not HDATF publications or measured product results.