AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning
- Published
- Source
- arXiv
- Paper number
- 833
- Field
- AI / General
- arXiv ID
- 2608.05987
Key points
- It aligns token-level distillation gaps with environment transition points by aggregating them at the turn level.
- It recursively updates a Bayesian belief state to compute credit that reflects the accumulated context of earlier turns.
- On ALFWorld with Qwen2.5-7B, the method achieves an 89.1% success rate, outperforming GRPO and existing self-distillation methods.
- The performance gap widens for longer horizons, while the effect is smaller on short tasks.
- It is fully compatible with standard policy optimization without extra rollouts or a learned critic.
Paper links
External research summaries. These are not HDATF publications or measured product results.