Agentic Memory: Learning Unified Long-Term and Short-Term Memory Management for Large Language Model Agents

Published
Source
arXiv
Paper number
108
Field
Memory / Agents
arXiv ID
2601.01885

Key points

  • LLM agents struggle with long-horizon reasoning because of their finite context window, which makes effective memory management necessary.
  • Traditional memory systems treat long-term and short-term memory as independent components, which leads to fragmented and heuristic-based solutions with suboptimal coordination.
  • Optimizing memory operations with standard reinforcement learning is difficult because reward signals are sparse and discontinuous over long trajectories.
  • AgeMem formalizes LTM operations, namely ADD, UPDATE, and DELETE, and STM operations, namely RETRIEVE, SUMMARY, and FILTER, as explicit tool-based actions within the agent policy, enabling autonomous and learnable memory control.
  • To learn comprehensive memory capability, it uses a three-stage progressive reinforcement learning strategy covering LTM construction, STM control with distractors, and integrated reasoning.
  • Stage-wise Group Relative Policy Optimization, or GRPO, ensures a consistent learning signal for memory decisions by propagating the terminal reward uniformly across trajectory steps, which addresses the sparse-reward problem.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)