JEV-as-a-Judge: Accept When Confident, Escalate When Unsure

Published
Source
arXiv
Paper number
1122
Field
AI Agents
arXiv ID
2609.26550

Key points

  • Compared a cheap verdict-plus-probability judge (JEV) fairly against sixteen generative LLM judges
  • Stayed within three percentage points of the strongest judge on ordinary preference and evidence-grounded factuality tasks
  • Measured cost was only 0.36% of the strongest comparator ($0.042 per million input tokens)
  • A cascade that accepts confident verdicts and escalates uncertain ones retained 99% of the comparator's accuracy
  • Showed clear limits on hard judgments such as checking derivations or resisting well-written wrong answers

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)