Ask, Don't Judge: Binary Questions for Interpretable LLM Evaluation and Self-Improvement
- Published
- Source
- arXiv
- Paper number
- 510
- Field
- AI / General
- arXiv ID
- 2606.27226
Key points
- It uses a meta-prompt structure that breaks evaluation criteria into atomic binary questions, which greatly improves transparency and diagnostic value.
- On QAGS for factual consistency, it improves Spearman rho by 0.195 over UniEval, which shows that the decomposition itself contributes most of the information gain.
- BINEVAL catches subtle factual errors that G-Eval and UniEval give a perfect 5.0 to by asking 7 questions each, reaching 1.57 points and coming close to human judgment at 2.0.
- It avoids the usual ceiling effect of LLM judges and distinguishes borderline outputs from clearly defective ones more effectively.
- It improves both summarization and IFBench prompts through cross-model prompt updates that use question-level feedback.
- Its task-agnostic and training-free design makes it immediately applicable to summarization, dialogue, instruction following, and other tasks.
Paper links
External research summaries. These are not HDATF publications or measured product results.