Self-Evaluation Is Already There: Eliciting Latent Judge Calibration in Base LLMs with Minimal Data
- Published
- Source
- arXiv
- Paper number
- 324
- Field
- LLMs / NLP
- arXiv ID
- 2606.05122
Key points
- This paper finds that the capability is generally present even before any targeted training. With only a few-shot prompt, the base model already predicts multi-attribute quality scores from an external judge on open-ended responses across three benchmarks far above chance.
- From just 160 unique examples, about 31 times fewer than the RL baseline, SEE improves held-out calibration across three benchmarks while preserving answer quality.
- These results recast judge-aligned self-evaluation as a problem of extraction rather than acquisition.
Paper links
External research summaries. These are not HDATF publications or measured product results.