Memory Layers at Scale

Published
Source
arXiv
Paper number
008
Field
Architecture
arXiv ID
2412.09764

Key points

  • Each memory layer acts as a low-cost learned lookup table that retrieves the top-k trainable keys for an input query and combines their values, rather than using attention over sequence activations.
  • Multiplicative keys reduce retrieval cost, values are sharded across GPUs, and custom CUDA kernels reach near-H100 memory bandwidth, enabling up to 128B-parameter memory in the reported scaling study.
  • As a result, Memory+ outperforms dense models with more than twice the compute budget and also beats MoE and PEER when compute and parameter counts are matched, with the clearest gains on factual memory benchmarks.
  • The implication is that the best design balances dense reasoning capacity with sparse memory capacity, which suggests that future LLMs can separate compute-intensive transformation from efficient knowledge storage.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)