$δ$-mem: Efficient Online Memory for Large Language Models

Published
Source
arXiv
Paper number
182
Field
LLMs / Memory / Architecture
arXiv ID
2605.12357

Key points

  • Large language models, or LLMs, struggle to use long interaction histories efficiently because of the quadratic compute cost of expanded context windows and context degradation.
  • Current memory mechanisms, including text-based, external-channel, and parametric approaches, suffer from information loss, retrieval noise, integration overhead, or the inability to adapt dynamically.
  • There is a need for a way to provide LLMs with continuous and adaptable memory without full model fine-tuning or large architectural changes.
  • delta-mem compresses continuous past information into a fixed-size associative memory online state, or OSAM, which enables efficient and dynamic storage.
  • The method operates through a sequential read-adjust-write cycle, where signals read from OSAM generate low-rank corrections to the frozen backbone's attention mechanism and a gated delta rule updates OSAM.
  • The system provides multiple write units, namely Token-State, Sequence-State, and Multi-State Write, so that memory updates can be adapted to different contexts and backbone capacities.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)