Cacheable by Design? Training Mixture-of-Experts Routers for Locality Against the Edge Memory-Bandwidth Wall: A Pre-Registered Negative Result with a Systems Measurement Study
arXiv cs.AIen
arXiv:2608.18261v1 Announce Type: new Abstract: Serving a 235B-parameter Mixture-of-Experts (MoE) model on a single 8 GB GPU is bottlenecked not by compute but by memory bandwidth: decode must stream each token's active experts from whichever tier holds them, and on consumer hardware most experts sit on an SSD far slower than RAM. We quantify this bandwidth wall on Qwen3-235B (Q4_K_M, 134 GB): measured decode is 0.44 tok/s warm, matching a bytes-per-token / bandwidth model, while a batching scheme that should amortize one disk sweep instead collapses at batch 32 from paging thrash. We build llama-moe-trace, a zero-surgery router-telemetry tool, and measure routing on Qwen3-30B: adjacent-toke
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Meta
- Verktyg
- Forskning
Related AI news
- How Unitree's Go series, which helped the company dominate the quadruped robot market, drew on openly published US university research funded by the US military (Michael Martina/Reuters)Techmeme · August 20, 2026
- Optimized Fuzzy Logic Approach with the IEEE Key Gas Method for Diagnosing Power Transformer Faults Using Dissolved Gas AnalysisarXiv cs.AI · August 20, 2026
- Position: Profiling Game Worlds by Transition ComplexityarXiv cs.AI · August 20, 2026
- SESSE: Sketch, Expand, Sort, Summarize, Evaluate -- LLM-as-Judge Evaluation via Structured DecompositionarXiv cs.AI · August 20, 2026
- Measuring the Partial-Credit Gap: A Strict Benchmark on Vietnam's 2025 Convex Marking SchemearXiv cs.AI · August 20, 2026
- FinSkillBench: Evaluating AI Agents and Domain Skills for Investment ManagementarXiv cs.AI · August 20, 2026