SAEM: Stage-Aware Expert Management for Memory-Efficient MoE Inference in Chain-of-Thought Reasoning
arXiv cs.AIen
arXiv:2608.21614v1 Announce Type: new Abstract: Chain-of-thought (CoT) prompting improves LLM reasoning by decomposing complex problems into intermediate steps, but its sequential nature increases decoding latency and memory usage. Mixture-of-Experts (MoE) models scale capacity through sparse expert activation, yet their full expert weights often exceed GPU memory and require costly GPU-CPU transfers. Existing runtimes treat all tokens uniformly, overlooking a key structural property of CoT traces: consecutive reasoning stages exhibit coherent and predictable expert activation patterns. Ignoring this stage-level regularity leads to inefficient caching and unnecessary data movement. We propos
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Forskning
Related AI news
- RIACT: A Responsible AI System for Personalized Study Habit Tracking and Early Burnout Signal Detection in University StudentsarXiv cs.AI · August 25, 2026
- Retrieval-grounded robot program generation and simulation-based correction via Model Context ProtocolarXiv cs.AI · August 25, 2026
- Quantifying geographic domain shift to decouple the geospatial transferability of human mobility flow generation modelsarXiv cs.AI · August 25, 2026
- A Reproducible, License-Aware Distillation Recipe for CPUDeployable Safety ClassificationarXiv cs.AI · August 25, 2026
- Data-Driven Dynamic Algorithm Dispatch with Large Language ModelsarXiv cs.AI · August 25, 2026
- Generate in the Chart, Not on the Boundary: Function-Symbol Grounding for Hard Constraints in LTN-GANsarXiv cs.AI · August 25, 2026