AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning

Published
Source
arXiv
Paper number
833
Field
AI / General
arXiv ID
2608.05987

Key points

  • It aligns token-level distillation gaps with environment transition points by aggregating them at the turn level.
  • It recursively updates a Bayesian belief state to compute credit that reflects the accumulated context of earlier turns.
  • On ALFWorld with Qwen2.5-7B, the method achieves an 89.1% success rate, outperforming GRPO and existing self-distillation methods.
  • The performance gap widens for longer horizons, while the effect is smaller on short tasks.
  • It is fully compatible with standard policy optimization without extra rollouts or a learned critic.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)