LLM-as-a-Verifier: A General-Purpose Verification Framework

Published
Source
arXiv
Paper number
573
Field
AI / General
arXiv ID
2607.05391

Key points

  • It computes continuous scores as the expected value of the probability distribution over scoring-token logits, greatly reducing ties.
  • It shows three verification scaling axes, score granularity g, repeat evaluation k, and criterion decomposition c, each improve accuracy.
  • It reaches state of the art on Terminal-Bench V2 at 86.5%, SWE-Bench Verified at 78.2%, RoboRewardBench at 87.4%, and MedAgentBench at 73.3%.
  • The fine-grained verification score can also be used for task progress estimation and as a continuous RL reward signal.
  • In LIBERO, using SAC reward shaping improves sample efficiency by about 1.8x, and in MATH with GRPO by about 1.1x.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)