Talk Isn't Always Cheap: Understanding Failure Modes in Multi-Agent Debate

Published
Source
arXiv
Paper number
084
Field
Multi-Agent / Debate
arXiv ID
2509.05396

Key points

  • There is a common belief that multi-agent debate universally improves large language model performance on complex reasoning tasks.
  • There is still a limited understanding of the specific failure modes that arise when multiple LLMs interact, especially in heterogeneous groups with varied capabilities.
  • It is necessary to clarify when and why multi-agent debate can be harmful rather than beneficial.
  • The authors used a standard multi-agent debate framework in which LLMs iteratively revise their responses over multiple rounds based on peer input.
  • They conducted an empirical study with heterogeneous groups of three LLMs, GPT-4o-mini, LLaMA-3.1-8B-Instruct, and Mistral-7B-Instruct-v0.2, across three datasets, CommonSenseQA, MMLU, and GSM8K.
  • They analyzed failure patterns through sequential revision traces, social influence evaluation using the probability of answer flipping after peer agreement, and interventions designed to reduce sycophancy.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)