The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation
- Published
- Source
- arXiv
- Paper number
- 733
- Field
- LLMs / NLP
- arXiv ID
- 2607.24720
Key points
- Making the world model explicit as a chain of thought greatly improves long-horizon generalization, from 93.1 versus 0.3.
- Atomic skills alone do not generalize compositionally, but adding only 5% long-horizon data causes a sharp jump.
- Mixing in bad trajectories causes error accumulation and collapses long-horizon performance.
- On-policy distillation has a wider effective range than GRPO under long-horizon and low-quality conditions.
- Multi-teacher distillation converges on shared planning patterns, while conflicting patterns cause severe forgetting.
- The experiments use full-parameter supervised learning based on Qwen2.5-100M.
Paper links
External research summaries. These are not HDATF publications or measured product results.