Sparse Delta Memory: Scaling the State of Linear RNNs through Sparsity

Published
Source
arXiv
Paper number
587
Field
Machine Learning
arXiv ID
2607.07386

Key points

  • Sparse addressing with product-key memory expands the state memory about 4,000 times beyond GDN at the 8B scale.
  • With the same FLOPs and the same number of parameters, it achieves better training loss than both GDN and full attention.
  • On long-context recall in RULER, it improves from 20.0 percent for GDN to 31.2 percent for SDM at 128k on the 1.4B model, which is an absolute gain of 11.2 percentage points.
  • A learned initial state, or learned M0, stores pretrained knowledge and further improves common-knowledge and reasoning tasks.
  • At the 8B scale, decoding is 1.49 times slower than GDN but 6 times faster than full attention.
  • The read and write distributions change adaptively with capacity: writes become concentrated, while reads become more distributed.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)