Proteus: Incremental Memory Activation for Long-Context Sequence Modeling

Published
Source
arXiv
Paper number
922
Field
Machine Learning
arXiv ID
2608.16844

Key points

  • The base experiment splits memory into 16 blocks, allows reading and writing only to the currently active blocks, and preserves locked blocks until they are reopened.
  • With 1.3 billion parameters and 100 billion training tokens, it reduces Titans perplexity on WikiText from 15.36 to 14.94 and on LAMBADA from 13.18 to 13.03.
  • After training on FineWeb with an 8K context and evaluating at 16K on S-NIAH-3, Titans improves from 21.4 to 29.8 accuracy, and on S-NIAH-2 it improves from 69.4 to 74.2.
  • Across six LongBench tasks, the average score rises from 15.72 to 16.65 for Hope-Attention, from 13.05 to 13.23 for Comba, and from 13.80 to 14.15 for Titans.
  • The optimal activation order is not searched separately, and the extension to MLP blocks is only a proof-of-concept checked in one Hope-Attention structure.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)