PHRBench: A Behavioral Evaluation of Post-Hallucination Reasoning in LLMs
- Published
- Source
- arXiv
- Paper number
- 1188
- Field
- LLMs / NLP
- arXiv ID
- 2610.10455
Key points
- Final-answer correctness alone cannot distinguish compliance, avoidance, and correction, so they built a framework that independently classifies each response trajectory into three behaviors (Hallucination Compliance / Avoidance / Heuristic Correction).
- Hallucinated context reduced accuracy by 7.7 percentage points on average, and across several model families they found a scaling tension: larger models become more vulnerable while also correcting more often.
- Successful recovery (correction plus reaching the correct answer) is rare, and recovered trajectories show a signature of more frequent belief updates.
- Using only prompt-level structural and semantic features, a lightweight predictor anticipated successful recovery before generation, achieving AUROC 0.847 and accuracy 0.863.
Paper links
External research summaries. These are not HDATF publications or measured product results.