Trace as State: Reasoning Traces as Conditional States for Long-Context Transformers
- Published
- Source
- arXiv
- Paper number
- 1065
- Field
- LLMs / NLP
- arXiv ID
- 2609.02702
Key points
- Simply placing the reasoning trace from the first reading before a long document and having the model read it again outperformed the existing approach of appending the trace in 26 of 27 model × task × metric combinations.
- A theoretical analysis explains the advantage of prepending the trace: providing a condition first requires exponentially less memory than providing it later.
- On GraphWalks Parents, GLM-5.2's accuracy rose from 66.4% on the first attempt to 100%, while DeepSeek V4 Pro improved from 29.2% to 81.8%.
- Inserting a trace from a different problem actually reduced performance. This indicates that the effect comes from information specific to the problem, rather than the format.
- Performance continued to improve as the number of traces increased from 1 to 5, making trace count an additional scaling lever for controlling the number of reasoning attempts.
- Its biggest advantage is that it requires no model-architecture changes and can therefore be added directly to any service that handles long documents.
Paper links
External research summaries. These are not HDATF publications or measured product results.