Decode-Latency Feedback Prefill: A Model-Free Controller and Its Generalization Limits
arXiv cs.AIen
arXiv:2609.38386v1 Announce Type: new Abstract: Concurrent autoregressive inference creates a fundamental interference problem: prefilling a newly arrived long prompt can delay tokens for requests that are already decoding. Fixed prefill chunks reduce this interference, but the best chunk size depends on the model, hardware, load, and latency objective. We introduce Decode-Latency Feedback Prefill (DLFP), a model-free controller that changes only prefill work that overlaps active decodes. After a guarded scheduling cycle, DLFP uses the observed interval as proportional feedback to resize the next prefill chunk; isolated prefills remain unrestricted. We implement DLFP in vLLM and evaluate it
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Forskning
Related AI news
- Kaikki noudattivat ohjeita – Kukaan ei ollut vastuussaTivi · October 1, 2026
- China's chip-tool localization accelerates: ACM Research backlog jumps 88%DIGITIMES · October 1, 2026
- ChartRevise: A Dataset and Evaluation Protocol for Exact Chart Editing via CodearXiv cs.AI · October 1, 2026
- Can an AI Agent Rediscover a Blaschke-Curve Invariant?arXiv cs.AI · October 1, 2026
- SimTrace: Grounded Multimodal User Trajectories Generation for Online User ModelingarXiv cs.AI · October 1, 2026
- ArgGYM: A Procedural, Engine-Verified Benchmark for Structured Defeasible ReasoningarXiv cs.AI · October 1, 2026