T^2MLR: Transformer with Temporal Middle-Layer Recurrence
- Published
- Source
- arXiv
- Paper number
- 650
- Field
- LLMs / NLP
- arXiv ID
- 2607.15178
Key points
- It shows that if only the middle layers, the transformer’s thinking layers, are connected recurrently, reasoning improves more than when the entire stack is repeated.
- Recursing through only 20% of the full stack is enough, which maximizes reasoning capability while minimizing inference latency.
- Even just adding the recurrent path and fine-tuning an already trained 1.7B model greatly improves mathematical reasoning accuracy, with GSM8K up 4.1 points and MATH500 up 5.2 points.
- The inference overhead is only about 8%, so it can be used without much burden in production.
- Training cost increases by roughly 2x to 4x, but it is a meaningful tradeoff given the gains in inference efficiency and performance.
Paper links
External research summaries. These are not HDATF publications or measured product results.