What Matters for Aggressive Decoding-Time KV Eviction? Temporal Aggregation and Ranking Preservation
arXiv cs.AIen
arXiv:2609.03515v1 Announce Type: new Abstract: Decoding-time KV cache compression research focuses heavily on designing better token scoring functions, while the temporal rule that aggregates scores across decode steps is often treated as an implementation detail. Under aggressive KV compression, we find that exponential-moving-average (EMA) aggregation makes approximately order-preserving scorer modifications largely indistinguishable at the eviction-set level. Value-norm and entropy variants remain highly correlated with attention and produce nearly unchanged retention sets, whereas KeyDiff, key norm, recency, and a learned scorer alter the ranking and degrade substantially. We associate
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Verktyg
- Forskning
Related AI news
- The sameness problem behind those unappetizing AI-generated menusTechCrunch AI · September 4, 2026
- Lite-On makes US$170 million strategic investment in DCX for AI cooling pushDIGITIMES · September 4, 2026
- K&S targets CPO, CoPoS growth with expanded TCB advanced packaging roadmapDIGITIMES · September 4, 2026
- Beyond "Made with AI": Visualizing Provenance Density to Mitigate the Transparency PenaltyarXiv cs.AI · September 4, 2026
- GPS-Bench: A Governance Policy Benchmark for Automating Policy AnalysisarXiv cs.AI · September 4, 2026
- HalluPeer: A Taxonomy-driven Benchmark for Detecting Hallucinations in Scientific Peer ReviewsarXiv cs.AI · September 4, 2026