Ask, Don't Judge: Binary Questions for Interpretable LLM Evaluation and Self-Improvement

Published
Source
arXiv
Paper number
510
Field
AI / General
arXiv ID
2606.27226

Key points

  • It uses a meta-prompt structure that breaks evaluation criteria into atomic binary questions, which greatly improves transparency and diagnostic value.
  • On QAGS for factual consistency, it improves Spearman rho by 0.195 over UniEval, which shows that the decomposition itself contributes most of the information gain.
  • BINEVAL catches subtle factual errors that G-Eval and UniEval give a perfect 5.0 to by asking 7 questions each, reaching 1.57 points and coming close to human judgment at 2.0.
  • It avoids the usual ceiling effect of LLM judges and distinguishes borderline outputs from clearly defective ones more effectively.
  • It improves both summarization and IFBench prompts through cross-model prompt updates that use question-level feedback.
  • Its task-agnostic and training-free design makes it immediately applicable to summarization, dialogue, instruction following, and other tasks.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)