KVBoost: Chunk-Level Key-Value Cache Reuse with Deviation-Guided Recomputation for Efficient Large Language Model Inference
arXiv cs.AIen
arXiv:2608.21362v1 Announce Type: new Abstract: Transformer-based large language models (LLMs) incur high prefill latency because key-value (KV) tensors must be recomputed for each request. Existing prefix-caching systems reduce this cost but require prompts to share a leading contiguous prefix, limiting effectiveness when shared content appears at arbitrary positions. We present KVBoost, a chunk-level KV cache reuse system for HuggingFace-compatible decoder models that enables reuse regardless of content position. KVBoost introduces a dual-hash keying scheme that separates positional identity (prefix hash) from content identity (content hash), supporting both exact and approximate cache mat
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Verktyg
- Forskning
Related AI news
- Désinformation russe : de fausses chaînes Telegram, de fausses vidéos et un vrai-faux think tank israélienLe Monde Pixels · August 25, 2026
- AI chipmaker Enflame sets subscription date for near $900 million Shanghai IPOEconomic Times Tech · August 25, 2026
- RIACT: A Responsible AI System for Personalized Study Habit Tracking and Early Burnout Signal Detection in University StudentsarXiv cs.AI · August 25, 2026
- Retrieval-grounded robot program generation and simulation-based correction via Model Context ProtocolarXiv cs.AI · August 25, 2026
- Quantifying geographic domain shift to decouple the geospatial transferability of human mobility flow generation modelsarXiv cs.AI · August 25, 2026
- A Reproducible, License-Aware Distillation Recipe for CPUDeployable Safety ClassificationarXiv cs.AI · August 25, 2026