Memory-R2: Fair Credit Assignment for Long-Horizon Memory-Augmented LLM Agents

Published
Source
arXiv
Paper number
213
Field
Machine Learning
arXiv ID
2605.21768

Key points

  • Surprisingly, these performance gains were achieved by training on only two full dialogue trajectories, which demonstrates the effectiveness of the RL paradigm.
  • It shows strong performance on out-of-distribution benchmarks such as LongMemEval and MSC-Self-Instruct.
  • The training gains observed on the 3B model transferred to the 7B model as well.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)