Agentic Memory: Learning Unified Long-Term and Short-Term Memory Management for Large Language Model Agents
- Published
- Source
- arXiv
- Paper number
- 108
- Field
- Memory / Agents
- arXiv ID
- 2601.01885
Key points
- LLM agents struggle with long-horizon reasoning because of their finite context window, which makes effective memory management necessary.
- Traditional memory systems treat long-term and short-term memory as independent components, which leads to fragmented and heuristic-based solutions with suboptimal coordination.
- Optimizing memory operations with standard reinforcement learning is difficult because reward signals are sparse and discontinuous over long trajectories.
- AgeMem formalizes LTM operations, namely ADD, UPDATE, and DELETE, and STM operations, namely RETRIEVE, SUMMARY, and FILTER, as explicit tool-based actions within the agent policy, enabling autonomous and learnable memory control.
- To learn comprehensive memory capability, it uses a three-stage progressive reinforcement learning strategy covering LTM construction, STM control with distractors, and integrated reasoning.
- Stage-wise Group Relative Policy Optimization, or GRPO, ensures a consistent learning signal for memory decisions by propagating the terminal reward uniformly across trajectory steps, which addresses the sparse-reward problem.
Paper links
External research summaries. These are not HDATF publications or measured product results.