ResearchPod Summary
Large language models (LLMs) are computationally expensive to deploy, largely due to the memory overhead of the key-value (KV) cache during long-context inference. While existing eviction strategies reduce memory usage by discarding less important tokens, they often suffer from significant performance degradation in long-context reasoning tasks. This paper investigates why these performance drops occur and proposes a more effective eviction strategy.
The authors identify that existing KV cache eviction methods suffer from a 'coverage problem': they tend to redundantly store similar tokens across different attention heads and layers, while failing to retain a diverse set of unique tokens. Using an information bottleneck framework, the authors theoretically demonstrate that reduced token coverage limits the mutual information between input and output, directly impairing predictive accuracy. To address this, they introduce K-VEC (KV Cache Eviction with Coverage), which adds two modules to the eviction process:
K-VEC consistently outperforms state-of-the-art eviction methods like SnapKV and PyramidKV across 16 subsets of the LongBench dataset. The improvements are most significant at lower cache budgets (e.g., 128 tokens), where existing methods struggle the most. The authors show that K-VEC maintains higher token coverage, which directly correlates with better F1 scores and task performance. While K-VEC incurs a minor increase in pre-fill latency, the overall decoding speed remains comparable to existing methods, making it a highly efficient solution for resource-constrained deployment.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.