Reasoning Structure of Large Language Models

Published
Source
arXiv
Paper number
354
Field
AI / General
arXiv ID
2606.03883

Key points

  • We build a scalable benchmark based on 21 two-dimensional grid logic puzzles, enabling controlled scaling studies across four difficulty levels.
  • We develop a pipeline that automatically converts unstructured text reasoning traces into atomic claim-dependency graphs.
  • We propose the reasoning efficiency metric eta, a topology-based measure that distinguishes focused reasoning from diffuse search.
  • We empirically show that token count is a poor proxy for reasoning quality, with a correlation of r = -0.05 with eta, and that extra tokens are mostly consumed by verification overhead.
  • Analysis of open-source reasoning models shows that the highest difficulty level is barely solved by any model, even with large token budgets, confirming the limits of simple compute scaling.
  • Early errors are strongly associated with inefficient traces, with r = 0.28, and eta increases when the time of error occurrence is later.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)