Can Agent Memory Systems Track Evolving State?

Published
Source
arXiv
Paper number
980
Field
AI / General
arXiv ID
2608.19652

Key points

  • It defined 'state drift': a failure in which an agent answers using old or invalidated values even when the required facts are present in retrieval results. It showed that this persists even under perfect retrieval, establishing it as a problem separate from 'retrieving well.'
  • It built StateMemBench with 234 multi-session scenarios and used closed-choice scoring to distinguish 'current state/old state/other failure,' cleanly isolating state-tracking failures.
  • StateMem explicitly tracks value replacement and dependencies (recomputation of derived values). It improved current-state accuracy 1.8-fold on DeepSeek-V4-Flash relative to the strongest same-backbone baseline (from 0.205 to 0.363), and 1.6-fold on Qwen-3.5-9B relative to the strongest existing memory system (from 0.149 to 0.233).
  • Even a single-call wrapper around existing memory systems improved all 6 backends by +32–+67 points. Comparisons with length- and cost-matched controls isolated +15–+32 points of this as the effect of the 'state structure' itself.
  • Judge labeling also confirmed that state drift accounts for a prominent share of failures on existing benchmarks such as LongMemEval.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)