Self-Evaluation Is Already There: Eliciting Latent Judge Calibration in Base LLMs with Minimal Data

Published
Source
arXiv
Paper number
324
Field
LLMs / NLP
arXiv ID
2606.05122

Key points

  • This paper finds that the capability is generally present even before any targeted training. With only a few-shot prompt, the base model already predicts multi-attribute quality scores from an external judge on open-ended responses across three benchmarks far above chance.
  • From just 160 unique examples, about 31 times fewer than the RL baseline, SEE improves held-out calibration across three benchmarks while preserving answer quality.
  • These results recast judge-aligned self-evaluation as a problem of extraction rather than acquisition.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)