Memory Layers at Scale
- Published
- Source
- arXiv
- Paper number
- 008
- Field
- Architecture
- arXiv ID
- 2412.09764
Key points
- Each memory layer acts as a low-cost learned lookup table that retrieves the top-k trainable keys for an input query and combines their values, rather than using attention over sequence activations.
- Multiplicative keys reduce retrieval cost, values are sharded across GPUs, and custom CUDA kernels reach near-H100 memory bandwidth, enabling up to 128B-parameter memory in the reported scaling study.
- As a result, Memory+ outperforms dense models with more than twice the compute budget and also beats MoE and PEER when compute and parameter counts are matched, with the clearest gains on factual memory benchmarks.
- The implication is that the best design balances dense reasoning capacity with sparse memory capacity, which suggests that future LLMs can separate compute-intensive transformation from efficient knowledge storage.
Paper links
External research summaries. These are not HDATF publications or measured product results.