Can I Buy Your KV Cache?

Published
Source
arXiv
Paper number
411
Field
AI / General
arXiv ID
2606.13361

Key points

  • KV cache reuse is token-exactly identical to prefill from scratch, with a 24 out of 24 greedy match and logits-level verification.
  • On Qwen3-4B, reused compute is 9 to 50 times cheaper than prefill, and the gap widens as documents get longer because of the L2 attention term.
  • The break-even point is effectively immediate, at N approximately 1, so reuse pays off from the second read onward.
  • KV is almost impossible to compress losslessly, so provider-side hosting is optimal and follows the same principle as production prompt caching.
  • Serving a 3,774-token document to 80M agents costs 0.5M with prefill versus 0.03M with reuse, a 49.7x reduction.
  • int8 quantization halves the size but breaks token-exactness, so it is unsuitable when losslessness is required.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)