Reinforcement Learning from Rich Feedback with Distributional DAgger

Published
Source
arXiv
Paper number
311
Field
Machine Learning
arXiv ID
2606.05152

Key points

  • The paper shows that RL methods using existing self-distillation objectives based on reverse KL or Jensen-Shannon do not guarantee monotonic policy improvement, meaning the update can increase the probability of worse actions even when the expert has higher reward.
  • In contrast, it shows that forward cross-entropy allows monotonic policy improvement and provides regret guarantees.
  • It also shows that this objective improves Pass@N by optimizing a lower bound on the teacher's weighted success likelihood.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)