Remember When It Matters: Proactive Memory Agent for Long-Horizon Agents

Published
Source
arXiv
Paper number
589
Field
AI / General
arXiv ID
2607.08716

Key points

  • It defines the concept of behavioral state decay, a failure mode in which information is present in context but still fails to affect behavior.
  • It uses a two-stage plug-and-play architecture in which a separate memory agent maintains a structured memory bank and decides whether to intervene selectively.
  • On Terminal-Bench 2.0, it improves Sonnet 4.5 from 37.6% to 45.9%, a gain of 8.3 points, and on τ²-Bench from 55.0% to 61.8%, a gain of 6.8 points.
  • The +2.4 to +2.5 point improvement also holds for the stronger Opus 4.6 model, so it remains effective even on strong models.
  • Catenation experiments show that active intervention outperforms passive memory exposure, always-on injection, advisor models, and Mem0 retrieval.
  • An open-weight memory agent trained from Qwen3.5-27B with SFT and GRPO also achieves a +3.5-point transfer gain on Terminal-Bench.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)