Calibrate, Then Route: A Measured Study of Learned Request Routing for Disaggregated LLM Serving
arXiv cs.AIen
arXiv cs.AI
AI Global WirearXiv:2609.16206v1 Announce Type: new Abstract: Disaggregated LLM serving places compute heavy prefill and memory heavy decode on separate GPU pools. Systems such as DistServe, Splitwise, and Mooncake make this separation fast, but routing still determines which instances handle each request. We study a router that estimates the additional completion time on each instance using exact prompt length, predicted output length, post admission KV cache pressure, and SLO class. We develop the policy in a discrete event simulator and validate it on eight NVIDIA A40 GPUs, each running a vLLM engine, with NIXL transferring KV caches between pools. All workloads run at measured saturation. Across three
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Forskning
- Reglering
Related AI news
- Pentagon CTO says U.S. government shouldn't take stakes in AI giants, questions adding AI rulesCNBC Technology · September 17, 2026
- Chip equipment and materials suppliers lead India investment pledges ahead of SEMICON India 2026DIGITIMES · September 17, 2026
- Elon Musk urges rival testing to expose AI safety flawsDIGITIMES · September 17, 2026
- House votes to curb AI data center costsAxios · September 16, 2026
- OpenAI CEO Sam Altman will attend state dinner for Trump-Xi summit in WashingtonCNBC Technology · September 16, 2026
- Anthropic and OpenAI want to embed safety evaluators. Will they really be independent?TechCrunch AI · September 16, 2026