Recirculation

Published
Source
arXiv
Paper number
944
Field
Machine Learning
arXiv ID
2608.17981

Key points

  • Transformers pass representations forward through their layers, so shallow layers cannot revise their interpretations using context-sensitive representations formed in deeper layers. Recirculation feeds those deeper representations back down, helping the model track state.
  • Without changing the model weights, and with only modest hyperparameter tuning of the adaptive transform, the method reduced perplexity by 23% and increased GSM8K accuracy by 21% across the Gemma 3 family.
  • The GSM8K error rate fell by 8.8% at pass@1 and by 20.9% at pass@128.
  • During generation, the two stacks can run in parallel with almost no additional cost. Prefilling, however, must process each token sequentially and may become expensive for long contexts.
  • The authors present this as an approach that draws clues for structural improvement from the properties of a trained model, rather than imposing an entirely new architecture.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)