LiveMCP-101: Stress Testing and Diagnosing MCP-enabled Agents on Challenging Queries

Published
Source
arXiv
Paper number
086
Field
Benchmarks / Agents
arXiv ID
2508.15760

Key points

  • AI agents show promising prototype performance, but they lack reliability and robustness when deployed in real, dynamic operating environments.
  • Existing benchmarks for tool-augmented LLMs are limited because they evaluate only single-step tool calls, operate in synthetic environments, or focus on low task complexity, so they miss the complexity of real tool interaction.
  • There has been no clear framework for diagnosing the concrete failure modes of agents that handle diverse, time-varying external tools.
  • We constructed LiveMCP-101, a benchmark of 101 carefully selected real-world queries that require multi-step coordination across diverse MCP-enabled tools.
  • We propose a new evaluation methodology that uses gold execution plans and parallel live execution to reflect the dynamic and evolving nature of live MCP tool responses.
  • We use a human-expert-validated LLM-as-a-judge approach to provide fine-grained scoring for both final outputs and the agent's execution trajectory.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)