Attention-Aware Routing: Coupling Routing and Attention in MoEs
arXiv cs.AIen
arXiv:2609.20974v1 Announce Type: new Abstract: In Mixture-of-Experts language models, the router typically selects and weights experts based on the token's hidden state, utilizing limited contextual information. We propose Attention-Aware Routing (AAR), which augments the router with temporal and spectral features extracted from a sliding window of attention weights that represent a summary of the model's contextual state, disentangled from the hidden state. Keeping the base transformer entirely frozen, we train only the routing parameters, isolating routing as the sole variable. AAR improves GSM8K by +3.37 pp over a routing-only SFT baseline on OLMoE. Beyond performance, we show that routi
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Forskning
Related AI news
- A researcher used GPT-6 Astra to decipher a WWI German radio transmission from 1918, one of the 50 famous unsolved ciphers listed on a German science blog (prinz)Techmeme · September 21, 2026
- DENSE: Distilling Agent Trajectories into Evidence-Grounded Shortcut Trees for Self-RefinementarXiv cs.AI · September 21, 2026
- GVPO++: Group Variance Policy Optimization for LLM Post-Training and On-Policy DistillationarXiv cs.AI · September 21, 2026
- Driving on Registers, Reasoning on Risk: Risk-Aware Occupancy for Register-Based End-to-End Autonomous DrivingarXiv cs.AI · September 21, 2026
- RBS-Attention: Radius-Bounded Sparse Prefill for Long-Context Large Language ModelsarXiv cs.AI · September 21, 2026
- Decoupling Internal Representational Changes and Causal Importance in Fine-Tuned Large Language ModelsarXiv cs.AI · September 21, 2026