SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning

Published
Source
arXiv
Paper number
639
Field
LLMs / NLP
arXiv ID
2607.14777

Key points

  • It creates a self-evolving loop that automatically extracts natural-language skills from completed trajectories and feeds them back into the policy.
  • It converts action-probability changes with and without a skill into a dense token-level supervision signal, OPD, and combines that with RL.
  • On ALFWorld, it reaches 91.8 percent, which is 7.4 points better than the previous best, and it beats GRPO trained on the full dataset even when using only 40 percent of the data.
  • It generalizes strongly to unseen environments, confirming that it learns real strategy rather than memorization.
  • It improves performance using only the trained policy, without extra prompts or external memory at inference time.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)