CompKV: Compensation-Aware KV Selection for Long-Context LLM Inference

Published
Source
arXiv
Paper number
1110
Field
Machine Learning
arXiv ID
2609.26300

Key points

  • Compensation-aware KV cache selection accounting for pruned token contributions
  • Accelerates inference with minimal degradation versus sparse attention
  • Relieves long-context inference bottlenecks
  • Applicable to cutting long-trajectory costs in agent harnesses

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)