What LLM Forecasters Know but Don't Say: Probing Internal Representations for Calibration and Faithfulness

Published
Source
arXiv
Paper number
588
Field
LLMs / NLP
arXiv ID
2607.08046

Key points

  • Internal probes cut ECE in half, from 0.093 to 0.044, compared with verbalized confidence, greatly improving calibration.
  • When CoT evidence is removed, predictions change while the reasoning trace stays the same in 23% of cases, revealing insincerity (rho = 0.22).
  • The probe's accuracy at tracking behavioral changes (rho = 0.57) far exceeds CoT (rho = 0.22), and it can detect 84% of the hidden effects with the correct direction.
  • Forced-answering experiments confirm that many predicted answers are already decided before reasoning even begins.
  • Using the entropy of the pre-reasoning answer distribution, we classify questions into three levels and save 30% to 47% of tokens without accuracy loss.
  • Probe-only calibration also works on GLM-4.7-Flash and GLM-4.5-Air, showing that calibration can improve without weight updates.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)