Recursive Experiential-Working Memory Evolution for Long-Horizon Agent Harnesses

Published
Source
arXiv
Paper number
1013
Field
AI / General
arXiv ID
2608.24876

Key points

  • It reduced noise from long conversation histories by retaining the verified current task state in working memory and using that state to retrieve needed skills from experiential memory.
  • It created a recursive improvement loop that uses execution traces to localize failures to memory components, modifies only the relevant part, and incorporates only changes that pass validation.
  • Of 37 completed combinations across four long-horizon tasks and ten models, 35 improved; even the best-performing models gained approximately 15–18 points, reaching a maximum success rate of 87.9%.
  • Gains increased with task length, reaching 32.2 points in the longest segment, and the largest reductions in representative long-horizon task failure types reached 80% or more.
  • However, failure localization is a decision about which part is most amenable to repair rather than a rigorous causal determination, and some verifier components were not sufficiently validated because they were ineffective on certain tasks.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)