RBS-Attention: Radius-Bounded Sparse Prefill for Long-Context Large Language Models
arXiv cs.AIen
arXiv:2609.20971v1 Announce Type: new Abstract: Long-context large language model inference is increasingly limited by prefill, where dense self-attention processes the entire prompt before generation begins. Sparse block selection can reduce this cost, but a block centroid may hide a highly relevant token among many irrelevant ones. We call this failure mode mean dilution and propose RBS-Attention, a training-free sparse-prefill method with two complementary selection branches. A centroid base branch captures average relevance, while a rescue branch uses the maximum key-block radius and its prompt-, layer-, and head-dependent distribution to identify blocks at risk of underestimation. Indep
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Forskning
Related AI news
- A researcher used GPT-6 Astra to decipher a WWI German radio transmission from 1918, one of the 50 famous unsolved ciphers listed on a German science blog (prinz)Techmeme · September 21, 2026
- DENSE: Distilling Agent Trajectories into Evidence-Grounded Shortcut Trees for Self-RefinementarXiv cs.AI · September 21, 2026
- GVPO++: Group Variance Policy Optimization for LLM Post-Training and On-Policy DistillationarXiv cs.AI · September 21, 2026
- Driving on Registers, Reasoning on Risk: Risk-Aware Occupancy for Register-Based End-to-End Autonomous DrivingarXiv cs.AI · September 21, 2026
- Decoupling Internal Representational Changes and Causal Importance in Fine-Tuned Large Language ModelsarXiv cs.AI · September 21, 2026
- Ability-Residual Decoupled Modeling for Affective Cognitive DiagnosisarXiv cs.AI · September 21, 2026