The Other Half of the Memory Wall: Serving 35B MoEs from SSD with Trained Routing Prediction
arXiv cs.AIen
arXiv:2609.18063v1 Announce Type: new Abstract: Mixture-of-experts (MoE) inference on consumer hardware is bounded by weight memory: a 35B-class model is 19.5GB at 4-bit, and sparsity shrinks the compute per token, not the bytes that must be held. Naive offloading to SSD does not help on its own, because layer N+1's experts must be chosen before layer N's output exists, so the reads cannot start early enough to hide behind compute. We present Edge0, a streaming MoE inference engine that closes the gap with a prerouter: a per-layer head predicts the next layer's routing one token ahead, and the prediction is consumed as the routing itself, so the staged expert set equals the routed set and no
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Forskning
Related AI news
- Sources: Emulate, a month-old UK AI startup founded by former Google DeepMind researchers, is in advanced talks to raise as much as $700M at a $3.7B valuation (Financial Times)Techmeme · September 17, 2026
- EvolveTrade: Experience-Driven Policy Refinement for Self-Evolving LLM Trading AgentsarXiv cs.AI · September 17, 2026
- SNOMED CT Concept Recommendation from Masked Clinical ContextarXiv cs.AI · September 17, 2026
- Memory Has Geometry: Non-Uniform Geometric Memory for Long-Horizon Personalized AIarXiv cs.AI · September 17, 2026
- When to Call an LLM: A Confidence-Gated Hybrid for Cost-Effective Emotion Recognition in Conversational AIarXiv cs.AI · September 17, 2026
- Multimodal Conditioning of Fine-Tuned Stable Diffusion XL for Controllable and Culturally Faithful Ulos Motif GenerationarXiv cs.AI · September 17, 2026