Continual Learning via Sparse Memory Finetuning
- Published
- Source
- arXiv
- Paper number
- 088
- Field
- Continual Learning
- arXiv ID
- 2510.15103
Key points
- Large language models suffer from catastrophic forgetting, meaning they lose previously learned knowledge when updated with new information.
- Current continual-learning methods for LLMs, such as replay-based and regularization-based approaches, are often compute-heavy, data-inefficient, or require trade-offs between new learning and retaining old knowledge.
- LLMs are usually deployed as static models after training, which limits their ability to adapt to real-time information, user feedback, and changing world knowledge.
- This paper proposes sparse memory fine-tuning, which uses a memory hierarchy where each token accesses only a small portion of a large memory pool.
- It applies TF-IDF-like ranking scores to identify the most knowledge-specific memory indices in a new training batch, reducing interference with general knowledge.
- During fine-tuning, only the top t memory indices identified by TF-IDF scores are updated, while all other memory parameters and the base LLM parameters remain fixed through gradient masking.
Paper links
External research summaries. These are not HDATF publications or measured product results.