Sample-Efficient Learning from Agent Experience

Published
Source
arXiv
Paper number
716
Field
LLMs / NLP
arXiv ID
2607.21051

Key points

  • The paper defines the problem of Experience Distillation, which distills an agent's interaction experience into weights, and proposes a solution.
  • It removes extra environment use by using a one-step branch rollout that resamples only the teacher's next action at each point in the collected trajectory.
  • It preserves 64.8 percent of the in-context learning effect, which greatly outperforms standard SFT at 3.8 percent recovery.
  • It matches reinforcement-learning baselines while using more than 9.6 times fewer environment samples.
  • The experimental details, including experience preprocessing, strengthened teacher inference, and branch packing, are validated empirically.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)