PHRBench: A Behavioral Evaluation of Post-Hallucination Reasoning in LLMs

Published
Source
arXiv
Paper number
1188
Field
LLMs / NLP
arXiv ID
2610.10455

Key points

  • Final-answer correctness alone cannot distinguish compliance, avoidance, and correction, so they built a framework that independently classifies each response trajectory into three behaviors (Hallucination Compliance / Avoidance / Heuristic Correction).
  • Hallucinated context reduced accuracy by 7.7 percentage points on average, and across several model families they found a scaling tension: larger models become more vulnerable while also correcting more often.
  • Successful recovery (correction plus reaching the correct answer) is rare, and recovered trajectories show a signature of more frequent belief updates.
  • Using only prompt-level structural and semantic features, a lightweight predictor anticipated successful recovery before generation, achieving AUROC 0.847 and accuracy 0.863.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)