OPID: On-Policy Skill Distillation for Agentic Reinforcement Learning
- Published
- Source
- arXiv
- Paper number
- 505
- Field
- LLMs / NLP
- arXiv ID
- 2606.26790
Key points
- Because it directly extracts hindsight skills from completed on-policy trajectories, no external skill library or retrieval is needed.
- It introduces a two-level hierarchy with episode-level skills for global workflows and failure-avoidance rules, and step-level skills for local decisions.
- Critical-first routing uses step-level skills at important moments and falls back to episode-level skills otherwise.
- It achieves 58.9 percent on ALFWorld, compared with 46.1 percent for GRPO, and 74.2 percent on WebShop, compared with 63.3 percent for GRPO.
- It provides a qualitative analysis showing that skill injection fixes the hallucinated-target error in GRPO-trained agents, where the agent searches for objects that do not exist.
- It is a training-time-only framework that needs no analyzer, external skill retrieval, or privileged context at inference time.
Paper links
External research summaries. These are not HDATF publications or measured product results.