Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL

Published
Source
arXiv
Paper number
074
Field
Agents / Distillation
arXiv ID
2508.13167

Key points

  • Existing multi-agent systems, or MAS, suffer from high computational overhead, generalization issues, and a lack of end-to-end data-driven training because LLM backbones are not inherently trained for multi-turn multi-agent workflows.
  • Existing tool-integrated reasoning, or TIR, models are constrained by rigid think-act-observe workflows and do not natively support flexible multi-agent collaboration patterns.
  • Both MAS and TIR paradigms fail to make effective use of end-to-end training to learn complex, dynamic multi-agent, multi-tool orchestration within a single unified LLM.
  • The Chain-of-Agents, or CoA, paradigm integrates multi-agent collaboration inside the decoding process of a single LLM, dynamically activating role agents such as thinking, planning, and reflection as well as tool agents such as search, crawling, and code generation.
  • An agentic supervised fine-tuning, or SFT, framework distills expert-level multi-agent trajectories into CoA-compatible training data and uses progressive quality filtering to ensure high-quality, complex problem-solving patterns.
  • Agentic reinforcement learning, or RL, further refines the LLM's multi-tool orchestration policy through outcome-based rewards, strengthening strategic tool use and adaptive reasoning on demanding task-specific queries.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)