On the Reliability Limits of LLM-Based Multi-Agent Planning
- Published
- Source
- arXiv
- Paper number
- 135
- Field
- Multi-Agent / Theory
- arXiv ID
- 2603.26993
Key points
- It models agents, messages, private tool signals, end actions, and optional human escalation within a single Bayesian risk framework rather than as role names such as planner or critic.
- It proves an upper bound on reliability: when there is no new exogenous information, additional agent steps can only transform or compress shared evidence, not beat a centralized Bayesian decision-maker who has that evidence.
- It characterizes inter-agent prose as a lossy communication channel. The relevant design metric is posterior distortion, not the number of agents, the number of rounds, or transcript length.
- The experiments show a large relay penalty. gpt-4.1-mini drops from 90.7 percent single-agent accuracy to 41.2 percent with two relay agents, 43.5 percent with three, and 22.5 percent with five, while posterior-vector communication reaches 75.2 percent versus 58.1 percent for prose on a matched subset.
- Tool use matters only when it changes the information structure. Wikipedia access is mostly redundant on MMLU, but synthetic KB lookups raise accuracy from 24.3 percent without tools to 82.7 percent.
Paper links
External research summaries. These are not HDATF publications or measured product results.