DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression
- Published
- Source
- arXiv
- Paper number
- 1094
- Field
- LLMs / NLP
- arXiv ID
- 2609.19969
Key points
- Cut per-token memory to about 1/4 of the previous generation (890 bytes) by combining cross-layer KV sharing (CSA2) with FP4 low-precision storage.
- Reduced the persistent SSD cache to about 1/8 with SWA Bounded Replay, which benefits agents that reuse long conversations across sessions.
- The encoder-decoder structure (CED) halves active parameters during prefill (8B), making it cost-efficient for agent workloads with heavy input and short output.
- Supports contexts up to one million tokens, pretrained on 45T multimodal tokens and post-trained with large-scale agent task synthesis plus RL.
- Confirmed across 8 benchmarks that raising the reasoning-effort setting from 25 to 100 predictably scales response length by 2.0-3.1x while consistently improving performance.
- Model checkpoints are publicly released on Hugging Face for immediate use.
Paper links
External research summaries. These are not HDATF publications or measured product results.