Failed Reasoning Traces Tell You What Is Fixable (But Not by Reading Them)

Published
Source
arXiv
Paper number
314
Field
Machine Learning
arXiv ID
2606.05145

Key points

  • The paper argues that this throws away an important signal, because some failures come from unlucky sampling and can be helped by more rollouts, while other failures are structural and resist resampling regardless of budget.
  • The authors propose that failed reasoning traces encode recoverability structure, meaning an inference-time signature of which test-time intervention can rescue a given failure.
  • Three task-level trajectory features derived from the structure of available interventions recover this structure from the distributional signature of failed rollouts rather than from the text of the rollouts themselves.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)