Reasoning Structure of Large Language Models
- Published
- Source
- arXiv
- Paper number
- 354
- Field
- AI / General
- arXiv ID
- 2606.03883
Key points
- We build a scalable benchmark based on 21 two-dimensional grid logic puzzles, enabling controlled scaling studies across four difficulty levels.
- We develop a pipeline that automatically converts unstructured text reasoning traces into atomic claim-dependency graphs.
- We propose the reasoning efficiency metric eta, a topology-based measure that distinguishes focused reasoning from diffuse search.
- We empirically show that token count is a poor proxy for reasoning quality, with a correlation of r = -0.05 with eta, and that extra tokens are mostly consumed by verification overhead.
- Analysis of open-source reasoning models shows that the highest difficulty level is barely solved by any model, even with large token budgets, confirming the limits of simple compute scaling.
- Early errors are strongly associated with inefficient traces, with r = 0.28, and eta increases when the time of error occurrence is later.
Paper links
External research summaries. These are not HDATF publications or measured product results.