SMat-Attention: Structured Long-Context Sequence Modeling
arXiv cs.AIen
arXiv cs.AI
AI Global WirearXiv:2609.36062v1 Announce Type: new Abstract: Long-context sequence models face a fundamental tradeoff: softmax attention uses flexible token-level interactions at quadratic cost, whereas linear attention obtains linear-time training and constant-time decoding by compressing history into a fixed-size state. In this work, we ask whether we can connect these regimes through a tunable notion of structure. To this end, we introduce Structured Matrix Attention (SMat-Attention) via a family of causal masks with structured long-range routing whose row supports have VC-dimension $d$. In our construction, $d=1$ recovers the standard causal mask, and increasing $d$ permits richer subset-routing patt
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Forskning
Related AI news
- Chinese firms trail global peers on profits, but AI power boom offers bright spot: NatixisSCMP Tech · September 30, 2026
- More Features Are Not More Evidence: Limits of Training-Free Human Activity Recognition with JevarXiv cs.AI · September 30, 2026
- Towards Mitigating Deceptive Safety Alignment in Large Reasoning ModelsarXiv cs.AI · September 30, 2026
- GeoOutageBench: Benchmarking Ambiguity-aware, Ontology-grounded Geospatiotemporal KGQA for Multimodal Power Outage and Resilience AnalysisarXiv cs.AI · September 30, 2026
- The Layer Mystery of VLA: An Information-Theoretical Analysis of VLA Latent InterfacearXiv cs.AI · September 30, 2026
- An Empirical Study and Assessment of EU AI Act Compliance CheckersarXiv cs.AI · September 30, 2026