The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation

Published
Source
arXiv
Paper number
733
Field
LLMs / NLP
arXiv ID
2607.24720

Key points

  • Making the world model explicit as a chain of thought greatly improves long-horizon generalization, from 93.1 versus 0.3.
  • Atomic skills alone do not generalize compositionally, but adding only 5% long-horizon data causes a sharp jump.
  • Mixing in bad trajectories causes error accumulation and collapses long-horizon performance.
  • On-policy distillation has a wider effective range than GRPO under long-horizon and low-quality conditions.
  • Multi-teacher distillation converges on shared planning patterns, while conflicting patterns cause severe forgetting.
  • The experiments use full-parameter supervised learning based on Qwen2.5-100M.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)