Benchmarking the Benchmarks: Evaluating Benchmarks for Conversational Agents
- Published
- Source
- arXiv
- Paper number
- 857
- Field
- LLMs / NLP
- arXiv ID
- 2608.06329
Key points
- It proposes an LLM-judge-based framework for evaluating the consistency, complexity, and policy coverage of conversational-agent benchmarks without reference labels.
- It validates the framework with benchmarks generated by LLMs of different strengths and consistently distinguishes quality differences across benchmarks made by models with different capabilities.
- The metrics also detect the expected drop in quality on deliberately degraded benchmarks.
- Correlation tests against human annotation show a medium-to-strong statistically significant correlation.
- It also applies the framework to manually curated benchmarks, such as tau3-bench, to demonstrate practical usefulness.
Paper links
External research summaries. These are not HDATF publications or measured product results.