TileMix: Tile-Centric Mixed-Precision Attention for LLM Inference Acceleration
arXiv cs.AIen
arXiv cs.AI
AI Global WirearXiv:2608.17336v1 Announce Type: new Abstract: Long-context prefill in large language models (LLMs) incurs substantial computation and memory traffic because dense self-attention computes quadratic query-key scores. Existing methods either use a uniform low-precision path or select token interactions, leaving spatial precision routing over hardware-aligned score tiles outside fused dense attention. We introduce TileMix, a tile-centric precision-routing kernel that makes numerical precision an executable spatial decision over score-tile groups within fused dense attention. TileMix partitions the attention matrix into hardware-aligned score tiles, packs routing decisions into compact bitmasks
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Forskning
Related AI news
- Depth Enables Local Entropy: Quadratic Depth Dependence in Deep Variation-Norm ReLU RegressionarXiv cs.AI · August 19, 2026
- Do LLMs Know a Good Hypothesis When They See One? Logit-Based Energy Scoring Outperforms Prompted LLM-as-Judge for Scientific Hypothesis RankingarXiv cs.AI · August 19, 2026
- KernelArc: A Multi-Agent Framework for GPU Kernel OptimizationarXiv cs.AI · August 19, 2026
- Synthesizing Feature Extractors: An Agentic Approach for Algorithm SelectionarXiv cs.AI · August 19, 2026
- PlanPO: Group Planning-Aware Policy Optimization for Multi-Turn Agentic LLMsarXiv cs.AI · August 19, 2026
- SkillEffect: Checked Lowering for Memory-Bounded Agent ToolsarXiv cs.AI · August 19, 2026