T^2MLR: Transformer with Temporal Middle-Layer Recurrence

Published
Source
arXiv
Paper number
650
Field
LLMs / NLP
arXiv ID
2607.15178

Key points

  • It shows that if only the middle layers, the transformer’s thinking layers, are connected recurrently, reasoning improves more than when the entire stack is repeated.
  • Recursing through only 20% of the full stack is enough, which maximizes reasoning capability while minimizing inference latency.
  • Even just adding the recurrent path and fine-tuning an already trained 1.7B model greatly improves mathematical reasoning accuracy, with GSM8K up 4.1 points and MATH500 up 5.2 points.
  • The inference overhead is only about 8%, so it can be used without much burden in production.
  • Training cost increases by roughly 2x to 4x, but it is a meaningful tradeoff given the gains in inference efficiency and performance.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)