OPID: On-Policy Skill Distillation for Agentic Reinforcement Learning

Published
Source
arXiv
Paper number
505
Field
LLMs / NLP
arXiv ID
2606.26790

Key points

  • Because it directly extracts hindsight skills from completed on-policy trajectories, no external skill library or retrieval is needed.
  • It introduces a two-level hierarchy with episode-level skills for global workflows and failure-avoidance rules, and step-level skills for local decisions.
  • Critical-first routing uses step-level skills at important moments and falls back to episode-level skills otherwise.
  • It achieves 58.9 percent on ALFWorld, compared with 46.1 percent for GRPO, and 74.2 percent on WebShop, compared with 63.3 percent for GRPO.
  • It provides a qualitative analysis showing that skill injection fixes the hallucinated-target error in GRPO-trained agents, where the agent searches for objects that do not exist.
  • It is a training-time-only framework that needs no analyzer, external skill retrieval, or privileged context at inference time.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)