Contiguity, Not Importance: Budgeted Repair of Stale KV Caches After Document Edits
arXiv cs.AIen
arXiv cs.AI
AI Global WirearXiv:2609.17983v1 Announce Type: new Abstract: KV-cache reuse can reduce inference cost in retrieval-augmented generation and agentic systems, but cached contexts may become stale when retrieved knowledge, working memory, or user state is edited. Under causal self-attention, even a local edit can affect downstream KV states. A full re-prefill reliably restores consistency but is costly, whereas refreshing only the edited span can leave downstream dependencies stale. We formulate in-place repair as budgeted recomputation and compare training-free position-selection policies on a factual RAG benchmark with matched direct and derived edits. Across three model families, all policies repair dire
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Forskning
- Agenter
Related AI news
- Open AI:n tekoälyagentit lähtivät jälleen laukalle – tekoäly-yhtiö paljasti kuusi uutta tapaustaYle Uutiset · September 17, 2026
- Sources: Emulate, a month-old UK AI startup founded by former Google DeepMind researchers, is in advanced talks to raise as much as $700M at a $3.7B valuation (Financial Times)Techmeme · September 17, 2026
- Last to board: why travel may be AI’s final frontierTech in Asia · September 17, 2026
- The Other Half of the Memory Wall: Serving 35B MoEs from SSD with Trained Routing PredictionarXiv cs.AI · September 17, 2026
- EvolveTrade: Experience-Driven Policy Refinement for Self-Evolving LLM Trading AgentsarXiv cs.AI · September 17, 2026
- SNOMED CT Concept Recommendation from Masked Clinical ContextarXiv cs.AI · September 17, 2026