FlashMemory-DeepSeek-V4: Lightning Index Ultra-Long Context via Lookahead Sparse Attention

Published
Source
arXiv
Paper number
373
Field
Machine Learning
arXiv ID
2606.09079

Key points

  • LSA is a paradigm that predicts future context demand and loads only the necessary KV chunks onto the GPU.
  • It decouples the indexer from the backbone, making it possible to train on a single H20 GPU in one hour.
  • It replaces fixed Top-k with dynamic chunk recall based on a sigmoid threshold.
  • It improves accuracy by 0.6% while using only 13.5% KV cache on average.
  • At 500K length, it cuts memory use by 90% and still improves accuracy by 1.9%.
  • It has a clear limitation: performance drops sharply on the dense-search MRCR benchmark, from 76% to 48%.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)