Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL
- Published
- Source
- arXiv
- Paper number
- 074
- Field
- Agents / Distillation
- arXiv ID
- 2508.13167
Key points
- Existing multi-agent systems, or MAS, suffer from high computational overhead, generalization issues, and a lack of end-to-end data-driven training because LLM backbones are not inherently trained for multi-turn multi-agent workflows.
- Existing tool-integrated reasoning, or TIR, models are constrained by rigid think-act-observe workflows and do not natively support flexible multi-agent collaboration patterns.
- Both MAS and TIR paradigms fail to make effective use of end-to-end training to learn complex, dynamic multi-agent, multi-tool orchestration within a single unified LLM.
- The Chain-of-Agents, or CoA, paradigm integrates multi-agent collaboration inside the decoding process of a single LLM, dynamically activating role agents such as thinking, planning, and reflection as well as tool agents such as search, crawling, and code generation.
- An agentic supervised fine-tuning, or SFT, framework distills expert-level multi-agent trajectories into CoA-compatible training data and uses progressive quality filtering to ensure high-quality, complex problem-solving patterns.
- Agentic reinforcement learning, or RL, further refines the LLM's multi-tool orchestration policy through outcome-based rewards, strengthening strategic tool use and adaptive reasoning on demanding task-specific queries.
Paper links
External research summaries. These are not HDATF publications or measured product results.