MemRL: Self-Evolving Agents via Runtime Reinforcement Learning on Episodic Memory

Published
Source
arXiv
Paper number
112
Field
Memory / RL / Agents
arXiv ID
2601.03192

Key points

  • AI agents struggle to continue acquiring skills after deployment without catastrophically forgetting previously learned knowledge.
  • Fine-tuning large language models for continual adaptation is computationally expensive and highly prone to catastrophic forgetting.
  • Existing retrieval-augmented generation systems are passive and often retrieve semantically relevant but low-utility noise from memory, which hinders learning from environmental feedback.
  • MemRL separates a stable, frozen LLM backbone from a dynamic, plastic external episodic memory, enabling continual learning without weight updates or catastrophic forgetting.
  • The system formulates memory retrieval as a memory-based Markov decision process, in which the retrieval policy, not the LLM parameters, is optimized by reinforcement learning.
  • It uses an intent-experience-utility triple-memory structure and a two-stage retrieval mechanism that combines semantic recall with value-aware selection to actively filter and exploit high-utility experiences.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)