Proteus: Incremental Memory Activation for Long-Context Sequence Modeling
- Published
- Source
- arXiv
- Paper number
- 922
- Field
- Machine Learning
- arXiv ID
- 2608.16844
Key points
- The base experiment splits memory into 16 blocks, allows reading and writing only to the currently active blocks, and preserves locked blocks until they are reopened.
- With 1.3 billion parameters and 100 billion training tokens, it reduces Titans perplexity on WikiText from 15.36 to 14.94 and on LAMBADA from 13.18 to 13.03.
- After training on FineWeb with an 8K context and evaluating at 16K on S-NIAH-3, Titans improves from 21.4 to 29.8 accuracy, and on S-NIAH-2 it improves from 69.4 to 74.2.
- Across six LongBench tasks, the average score rises from 15.72 to 16.65 for Hope-Attention, from 13.05 to 13.23 for Comba, and from 13.80 to 14.15 for Titans.
- The optimal activation order is not searched separately, and the extension to MLP blocks is only a proof-of-concept checked in one Hope-Attention structure.
Paper links
External research summaries. These are not HDATF publications or measured product results.