Hallucinations Undermine Trust; Metacognition is a Way Forward
- Published
- Source
- arXiv
- Paper number
- 215
- Field
- AI Safety / Trust
- arXiv ID
- 2605.01428
Key points
- The central argument is that most factuality gains come from expanding the boundary of knowledge, not from improving awareness of that boundary, meaning the distinction between what the model knows and does not know.
- The model may fundamentally lack the ability to separate truth from error, which creates an inherent tradeoff between removing hallucinations and preserving usefulness.
- The paper proposes faithful uncertainty, meaning that the model's linguistic uncertainty should match its intrinsic uncertainty, which offers a third path beyond either answering or refusing.
- For agent systems, uncertainty awareness becomes a control layer that governs when to explore, what to trust, and how to weight retrieved information.
- Empirical evidence, including poor generalization of truth probes, confident hallucinations, and failures of alignment methods, supports the existence of the discriminative gap.
Paper links
External research summaries. These are not HDATF publications or measured product results.