Are We Ready For An Agent-Native Memory System?
- Published
- Source
- arXiv
- Paper number
- 499
- Field
- LLMs / NLP
- arXiv ID
- 2606.24775
Key points
- We propose an analytical framework that systematically decomposes agent memory into four modules: representation and storage, extraction, retrieval and routing, and maintenance.
- We evaluate 12 memory systems plus two baselines on a unified testbed spanning five workloads and 11 datasets.
- No single architecture is universally best. Hybrid systems are strong for conversational QA and graph-based systems are strong for factual recall, but they are weak at temporal reasoning.
- Memory extraction should preserve context at write time. Broader retention, rather than overly selective extraction, is better for downstream reasoning.
- Conservative merging outperforms delayed flushing or excessive summarization in long-term stability.
- Highly structured systems pay tens to hundreds of times more in index-build time and query latency, but the accuracy gains do not scale proportionally.
Paper links
External research summaries. These are not HDATF publications or measured product results.