Single-Agent LLMs Outperform Multi-Agent Systems on Multi-Hop Reasoning Under Equal Thinking Token Budgets

Published
Source
arXiv
Paper number
138
Field
Agents / Multi-Agent
arXiv ID
2604.02460

Key points

  • It is unclear whether the performance gains in multi-agent LLM systems are due to structural advantages or simply to higher test-time compute.
  • Prior comparisons between MAS and single-agent systems are often confounded by unequal token usage, which makes direct comparisons unreliable.
  • A robust evaluation method that strictly controls thought-token budgets is needed to clarify the source of the gains.
  • They introduce a theoretical framework based on the Data Processing Inequality to argue for the information efficiency of single-agent systems under a fixed reasoning budget.
  • They extensively evaluate a range of single-agent and multi-agent architectures, including Sequential, Debate, and Ensemble, across three LLM families, Qwen3, DeepSeek, and Gemini, on the multi-hop reasoning tasks FRAMES and MuSiQue.
  • They enforce strict control of the thought-token budget, meaning intermediate reasoning tokens excluding the prompt and final answer, to ensure fair comparisons across systems.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)