Random Attention: Rethinking KV Cache Eviction for Efficient Reasoning

Published
Source
arXiv
Paper number
1071
Field
LLMs / NLP
arXiv ID
2609.03430

Key points

  • The method preserves the input question while evicting cache entries uniformly at random from each attention head's reasoning history, eliminating importance scoring and ranking.
  • Comparisons across four models and six reasoning tasks showed accuracy comparable to strong existing eviction methods, and most performance differences shrank when all methods were configured to preserve the input question.
  • In vLLM experiments on a single H200 with a cache budget of 2048, a 1k-token input, and a 32k-token generation, output throughput was 32–43% higher than TriAttention.
  • The analysis that needed information is restated during reasoning and stored redundantly across multiple heads suggests that memory demands and cache-management costs in long-reasoning services could be reduced without complex importance calculations.
  • Selection signals helped in an experiment retrieving information that appeared only once and was never mentioned again, so these findings should not be generalized to mean that random eviction is superior for all information-retrieval tasks or all models.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)