Where Should the KV Cache Live? Placement Policies Across GPU, CPU, and SSD for Long-Lived Sessions
arXiv cs.AIen
arXiv cs.AI
AI Global WirearXiv:2609.16215v1 Announce Type: new Abstract: GPU high bandwidth memory is scarce and expensive, and KV caches consume much of it as chats, agent loops, and document question answering accumulate state. Systems such as Mooncake, LMCache, FlexGen, InfiniGen, and AttentionStore extend GPU memory with CPU DRAM and SSD. The harder question is which blocks belong in each tier, when to move or evict them, and whether prefetching helps. We study these choices in a discrete event simulator spanning GPU HBM, CPU DRAM, and SSD, calibrated against a random forest execution time predictor. We compare recency, reuse frequency, predicted reuse, and an EWMA predictor with prefetch lookahead across chat,
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Forskning
- Agenter
Related AI news
- Huawei sees AI agents driving 90% of token traffic by 2035DIGITIMES · September 17, 2026
- OpenAI unveils new framework for reporting ‘AI misalignment’ as it reveals six more worrying incidentsSiliconANGLE · September 17, 2026
- Chip equipment and materials suppliers lead India investment pledges ahead of SEMICON India 2026DIGITIMES · September 17, 2026
- AI 晶片愈做愈大,工程師卻不夠用了!Cadence 把 Agent 帶進 EDA 核心流程TechNews (TW) · September 17, 2026
- Agents – not humans – have picked the next $100b companyTech in Asia · September 17, 2026
- Elon Musk urges rival testing to expose AI safety flawsDIGITIMES · September 17, 2026