Nexus: Depth-Adaptive KV-Cache Splicing and Retrieval-Decoupled Tool Routing for Agentic LLMs on Unified Memory
arXiv cs.AIen
arXiv cs.AI
AI Global WirearXiv:2608.20397v1 Announce Type: new Abstract: Agentic large language models (LLMs) on the Model Context Protocol (MCP) re-encode verbose tool schemas every turn, so prefill - quadratic in sequence length - dominates time-to-first-token (TTFT) as the tool registry grows. Nexus's primary lever is to decouple routing from the schema-prefill cost: an INT8 semantic lookaside buffer (SLB) with a calibrated cross-encoder margin gate selects tools by retrieval, and arguments are generated over a compressed textual signature (median 19 tokens) rather than over spliced key/value (KV) cache. This path is depth-independent: routing accuracy stays near 89% as the registry scales to 250 tools - where a
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Verktyg
- Forskning
- Agenter
Related AI news
- ChatGPT får nye funktioner og ændringer: En af dem er gigantisk bagdør til din iPhoneIngeniøren · August 24, 2026
- Source: AI researcher Luke Metz, who returned to OpenAI from TML earlier this year, joins Meta's Superintelligence Labs and will report to Alexandr Wang (Ina Fried/Axios)Techmeme · August 24, 2026
- Analysis: SK Hynix pushes beyond HBM with HBF and CPODIGITIMES · August 24, 2026
- Dual-Cache Latent Space Communication between Heterogeneous Language ModelsarXiv cs.AI · August 24, 2026
- SDAD: Spec-Driven Agentic Development for the AI-Native SDLCarXiv cs.AI · August 24, 2026
- Who Delegates to AI? Evidence from 53,000 Agent ConfigurationsarXiv cs.AI · August 24, 2026