Memory Intelligence Agent

Published
Source
arXiv
Paper number
137
Field
Agents / Memory
arXiv ID
2604.04503

Key points

  • Existing memory systems for Deep Research Agents primarily rely on long-context memory, which leads to attention dilution, noise injection, and high storage and compute costs.
  • Most existing memory systems capture knowledge-oriented memory, or what, but fail to use process-oriented memory and conceptual knowledge, or how, which are central for guiding future planning in deep research.
  • The poor planning and execution of DRA come from the fact that, when memory is used only for few-shot Chain-of-Thought (CoT), the planner lacks task-specific learning and the executor misreads instructions.
  • MIA proposes a Manager-Planner-Executor architecture that separates nonparametric memory storage of past search trajectories from parametric planning and execution, and it uses a hybrid retrieval strategy.
  • A two-stage alternating reinforcement learning (RL) strategy promotes synergistic co-evolution between the Planner and the Executor, ensuring alignment between high-level planning and effective low-level task execution.
  • Test-Time Learning (TTL) and a new unsupervised evaluation framework enable continuous self-evolution, allowing the agent to refine both nonparametric and parametric memory online without explicit supervision.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)