Agent-as-a-Judge

Published
Source
arXiv
Paper number
110
Field
Evaluation / Agents
arXiv ID
2601.05111

Key points

  • Traditional AI evaluation methods, including predefined metrics and human judgment, have been insufficient to capture semantic nuance, or they have been hard to scale and expensive.
  • Early LLM-as-a-Judge paradigms face limitations such as intrinsic parameter bias, shallow single-pass reasoning, lack of real-world verification, and cognitive overload on complex, multi-dimensional tasks.
  • The rapidly growing field of Agent-as-a-Judge lacked a unified framework, comprehensive taxonomy, and clear roadmap for organizing its diverse methods and applications.
  • This paper introduces a new three-stage maturity taxonomy for agentic evaluation systems based on autonomy and adaptability: Procedural, Reactive, and Self-Evolving Agent-as-a-Judge.
  • It systematically classifies core agentic methodologies, including multi-agent collaboration, planning, tool integration, memory and personalization, and optimization paradigms.
  • The survey comprehensively reviews Agent-as-a-Judge applications across broad general domains, such as mathematics, code, and fact checking, as well as specialized domains, such as medicine, law, finance, and education.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)