FlashMemory-DeepSeek-V4: Lightning Index Ultra-Long Context via Lookahead Sparse Attention
- Published
- Source
- arXiv
- Paper number
- 373
- Field
- Machine Learning
- arXiv ID
- 2606.09079
Key points
- LSA is a paradigm that predicts future context demand and loads only the necessary KV chunks onto the GPU.
- It decouples the indexer from the backbone, making it possible to train on a single H20 GPU in one hour.
- It replaces fixed Top-k with dynamic chunk recall based on a sigmoid threshold.
- It improves accuracy by 0.6% while using only 13.5% KV cache on average.
- At 500K length, it cuts memory use by 90% and still improves accuracy by 1.9%.
- It has a clear limitation: performance drops sharply on the dense-search MRCR benchmark, from 76% to 48%.
Paper links
External research summaries. These are not HDATF publications or measured product results.