Memory Attention

Published
Source
arXiv
Paper number
1111
Field
Machine Learning
arXiv ID
2609.28399

Key points

  • Eliminates the dedicated value projection and replaces it with V = K + Norm(M), reducing value construction to a table lookup and an addition.
  • The number of training tokens needed to reach the same loss dropped by 29.6% for the smaller model and 13.8% for the larger one (token efficiency of 1.42x and 1.16x).
  • MA-Offload stores the memory tables in CPU memory, cutting GPU parameter storage by 55.45% while prefetching keeps inference speed nearly identical to the baseline (prefill 90.69ms vs 90.97ms).
  • Total parameters grew to about 2.08x the baseline, but because the extra memory is lookup-based, model capacity increased without adding dense computation.
  • Gains were also confirmed on retrieval tasks at up to twice the training context length.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)