Sparse Delta Memory: Scaling the State of Linear RNNs through Sparsity
- Published
- Source
- arXiv
- Paper number
- 587
- Field
- Machine Learning
- arXiv ID
- 2607.07386
Key points
- Sparse addressing with product-key memory expands the state memory about 4,000 times beyond GDN at the 8B scale.
- With the same FLOPs and the same number of parameters, it achieves better training loss than both GDN and full attention.
- On long-context recall in RULER, it improves from 20.0 percent for GDN to 31.2 percent for SDM at 128k on the 1.4B model, which is an absolute gain of 11.2 percentage points.
- A learned initial state, or learned M0, stores pretrained knowledge and further improves common-knowledge and reasoning tasks.
- At the 8B scale, decoding is 1.49 times slower than GDN but 6 times faster than full attention.
- The read and write distributions change adaptively with capacity: writes become concentrated, while reads become more distributed.
Paper links
External research summaries. These are not HDATF publications or measured product results.