DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression

Published
Source
arXiv
Paper number
1094
Field
LLMs / NLP
arXiv ID
2609.19969

Key points

  • Cut per-token memory to about 1/4 of the previous generation (890 bytes) by combining cross-layer KV sharing (CSA2) with FP4 low-precision storage.
  • Reduced the persistent SSD cache to about 1/8 with SWA Bounded Replay, which benefits agents that reuse long conversations across sessions.
  • The encoder-decoder structure (CED) halves active parameters during prefill (8B), making it cost-efficient for agent workloads with heavy input and short output.
  • Supports contexts up to one million tokens, pretrained on 45T multimodal tokens and post-trained with large-scale agent task synthesis plus RL.
  • Confirmed across 8 benchmarks that raising the reasoning-effort setting from 25 to 100 predictably scales response length by 2.0-3.1x while consistently improving performance.
  • Model checkpoints are publicly released on Hugging Face for immediate use.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)