The Illusion of $\textit{What If}$: Evaluating the Breakdown of Counterfactual Reasoning in LLMs

Published
Source
arXiv
Paper number
1051
Field
AI / General
arXiv ID
2608.27953

Key points

  • It released a diagnostic benchmark of 220 open-domain what-if questions spanning STEM, humanities and social sciences, and mixed domains.
  • It developed the PRISM evaluator, which converts free-form answers into causal graphs, DAGs of events, states, and mechanisms, to score the reasoning process rather than the outcome.
  • The best score among 6 current LLMs was only 64.62%, showing that counterfactual reasoning remains far from saturation.
  • The most common failure was a causal gap: models listed plausible outcomes but omitted the intermediate links connecting the premise to those outcomes.
  • It concluded that fluent narratives conceal weak causal processes, emphasizing the need for structure-based evaluation.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)