MAS-PromptBench: When Does Prompt Optimization Improve Multi-Agent LLM Systems?

Published
Source
arXiv
Paper number
475
Field
Machine Learning
arXiv ID
2606.23664

Key points

  • A benchmark dedicated to prompt optimization for MAS, covering five topologies, three communication protocols, and team sizes from 2 to 10 agents.
  • The gains range dramatically from as much as +24 points to as much as -16 points, proving that blind application is risky.
  • Optimization works better when communication protocols are more structured: Freeform < Semi-structured < Structured.
  • As team size grows, coordination complexity increases and optimization gains shrink, from +2.4 points for 2 agents to -2.1 points on average for 10 agents.
  • The natural extension of single-agent optimization tools to MAS is not sufficient, and MAS-specific algorithms are needed.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)