JEV-as-a-Judge: Accept When Confident, Escalate When Unsure
- Published
- Source
- arXiv
- Paper number
- 1122
- Field
- AI Agents
- arXiv ID
- 2609.26550
Key points
- Compared a cheap verdict-plus-probability judge (JEV) fairly against sixteen generative LLM judges
- Stayed within three percentage points of the strongest judge on ordinary preference and evidence-grounded factuality tasks
- Measured cost was only 0.36% of the strongest comparator ($0.042 per million input tokens)
- A cascade that accepts confident verdicts and escalates uncertain ones retained 99% of the comparator's accuracy
- Showed clear limits on hard judgments such as checking derivations or resisting well-written wrong answers
Paper links
External research summaries. These are not HDATF publications or measured product results.