SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning
- Published
- Source
- arXiv
- Paper number
- 639
- Field
- LLMs / NLP
- arXiv ID
- 2607.14777
Key points
- It creates a self-evolving loop that automatically extracts natural-language skills from completed trajectories and feeds them back into the policy.
- It converts action-probability changes with and without a skill into a dense token-level supervision signal, OPD, and combines that with RL.
- On ALFWorld, it reaches 91.8 percent, which is 7.4 points better than the previous best, and it beats GRPO trained on the full dataset even when using only 40 percent of the data.
- It generalizes strongly to unseen environments, confirming that it learns real strategy rather than memorization.
- It improves performance using only the trained policy, without extra prompts or external memory at inference time.
Paper links
External research summaries. These are not HDATF publications or measured product results.