LiveMCP-101: Stress Testing and Diagnosing MCP-enabled Agents on Challenging Queries
- Published
- Source
- arXiv
- Paper number
- 086
- Field
- Benchmarks / Agents
- arXiv ID
- 2508.15760
Key points
- AI agents show promising prototype performance, but they lack reliability and robustness when deployed in real, dynamic operating environments.
- Existing benchmarks for tool-augmented LLMs are limited because they evaluate only single-step tool calls, operate in synthetic environments, or focus on low task complexity, so they miss the complexity of real tool interaction.
- There has been no clear framework for diagnosing the concrete failure modes of agents that handle diverse, time-varying external tools.
- We constructed LiveMCP-101, a benchmark of 101 carefully selected real-world queries that require multi-step coordination across diverse MCP-enabled tools.
- We propose a new evaluation methodology that uses gold execution plans and parallel live execution to reflect the dynamic and evolving nature of live MCP tool responses.
- We use a human-expert-validated LLM-as-a-judge approach to provide fine-grained scoring for both final outputs and the agent's execution trajectory.
Paper links
External research summaries. These are not HDATF publications or measured product results.