Memory Attention
- Published
- Source
- arXiv
- Paper number
- 1111
- Field
- Machine Learning
- arXiv ID
- 2609.28399
Key points
- Eliminates the dedicated value projection and replaces it with V = K + Norm(M), reducing value construction to a table lookup and an addition.
- The number of training tokens needed to reach the same loss dropped by 29.6% for the smaller model and 13.8% for the larger one (token efficiency of 1.42x and 1.16x).
- MA-Offload stores the memory tables in CPU memory, cutting GPU parameter storage by 55.45% while prefetching keeps inference speed nearly identical to the baseline (prefill 90.69ms vs 90.97ms).
- Total parameters grew to about 2.08x the baseline, but because the extra memory is lookup-based, model capacity increased without adding dense computation.
- Gains were also confirmed on retrieval tasks at up to twice the training context length.
Paper links
External research summaries. These are not HDATF publications or measured product results.