Memory-R2: Fair Credit Assignment for Long-Horizon Memory-Augmented LLM Agents
- Published
- Source
- arXiv
- Paper number
- 213
- Field
- Machine Learning
- arXiv ID
- 2605.21768
Key points
- Surprisingly, these performance gains were achieved by training on only two full dialogue trajectories, which demonstrates the effectiveness of the RL paradigm.
- It shows strong performance on out-of-distribution benchmarks such as LongMemEval and MSC-Self-Instruct.
- The training gains observed on the 3B model transferred to the 7B model as well.
Paper links
External research summaries. These are not HDATF publications or measured product results.