Continual Learning Bench: Evaluating Frontier AI Systems in Real-World Stateful Environments

Published
Source
arXiv
Paper number
362
Field
AI / General
arXiv ID
2606.05661

Key points

  • It is a set of high-quality, expert-validated tasks from six real-world domains, and each task contains latent structure that cannot be solved through pretraining alone.
  • It proposes an evaluation methodology that introduces a gain metric to separate base model capability from experience-based learning.
  • Simple in-context learning (ICL) outperforms dedicated memory systems such as Mem0 and ACE on most tasks.
  • Even the best system reaches only 25.4% normalized gain, suggesting that the continual learning ability of current frontier models is still limited.
  • Memory modules are frequently observed to hurt performance by producing false generalizations and stale beliefs.
  • In terms of cost efficiency, ICL is also superior, and expensive systems fail to deliver proportional gains.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)