Continual Learning via Sparse Memory Finetuning

Published
Source
arXiv
Paper number
088
Field
Continual Learning
arXiv ID
2510.15103

Key points

  • Large language models suffer from catastrophic forgetting, meaning they lose previously learned knowledge when updated with new information.
  • Current continual-learning methods for LLMs, such as replay-based and regularization-based approaches, are often compute-heavy, data-inefficient, or require trade-offs between new learning and retaining old knowledge.
  • LLMs are usually deployed as static models after training, which limits their ability to adapt to real-time information, user feedback, and changing world knowledge.
  • This paper proposes sparse memory fine-tuning, which uses a memory hierarchy where each token accesses only a small portion of a large memory pool.
  • It applies TF-IDF-like ranking scores to identify the most knowledge-specific memory indices in a new training batch, reducing interference with general knowledge.
  • During fine-tuning, only the top t memory indices identified by TF-IDF scores are updated, while all other memory parameters and the base LLM parameters remain fixed through gradient masking.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)