Pretraining Recurrent Networks without Recurrence

Published
Source
arXiv
Paper number
332
Field
Machine Learning
arXiv ID
2606.06479

Key points

  • It reduces RNN training to supervised learning over stepwise memory transitions and bypasses recurrent backpropagation.
  • It learns what to remember by generating memory labels with a Transformer-based encoder.
  • It captures long-range dependencies stably with an O(1) gradient path between every token pair.
  • It enables parallel training along the time axis, which has the potential to remove the bottleneck in scaling RNNs.
  • It outperforms standard BPTT on language modeling and pixel sequence modeling.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)