What LLM Forecasters Know but Don't Say: Probing Internal Representations for Calibration and Faithfulness
- Published
- Source
- arXiv
- Paper number
- 588
- Field
- LLMs / NLP
- arXiv ID
- 2607.08046
Key points
- Internal probes cut ECE in half, from 0.093 to 0.044, compared with verbalized confidence, greatly improving calibration.
- When CoT evidence is removed, predictions change while the reasoning trace stays the same in 23% of cases, revealing insincerity (rho = 0.22).
- The probe's accuracy at tracking behavioral changes (rho = 0.57) far exceeds CoT (rho = 0.22), and it can detect 84% of the hidden effects with the correct direction.
- Forced-answering experiments confirm that many predicted answers are already decided before reasoning even begins.
- Using the entropy of the pre-reasoning answer distribution, we classify questions into three levels and save 30% to 47% of tokens without accuracy loss.
- Probe-only calibration also works on GLM-4.7-Flash and GLM-4.5-Air, showing that calibration can improve without weight updates.
Paper links
External research summaries. These are not HDATF publications or measured product results.