Trace as State: Reasoning Traces as Conditional States for Long-Context Transformers

Published
Source
arXiv
Paper number
1065
Field
LLMs / NLP
arXiv ID
2609.02702

Key points

  • Simply placing the reasoning trace from the first reading before a long document and having the model read it again outperformed the existing approach of appending the trace in 26 of 27 model × task × metric combinations.
  • A theoretical analysis explains the advantage of prepending the trace: providing a condition first requires exponentially less memory than providing it later.
  • On GraphWalks Parents, GLM-5.2's accuracy rose from 66.4% on the first attempt to 100%, while DeepSeek V4 Pro improved from 29.2% to 81.8%.
  • Inserting a trace from a different problem actually reduced performance. This indicates that the effect comes from information specific to the problem, rather than the format.
  • Performance continued to improve as the number of traces increased from 1 to 5, making trace count an additional scaling lever for controlling the number of reasoning attempts.
  • Its biggest advantage is that it requires no model-architecture changes and can therefore be added directly to any service that handles long documents.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)