Reasoning-Aware Compression: Identifying and Protecting Vulnerable Reasoning Circuits for Energy-Efficient LLM Deployment
arXiv cs.AIen
arXiv:2609.05512v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) impose substantial energy costs during deployment, yet current compression methods apply uniform quantization across all components, risking damage to critical reasoning circuits. We present a reasoning-aware compression framework that benchmarks quantization conditions across five reasoning benchmarks, GSM8K, FOLIO, MATH-500, ProofWriter, and MuSiQue, with hardware-level GPU energy measurement; profiles per-module INT4 vulnerability across all 196-224 (layer, projection) pairs via a perturbation sweep on a held-out calibration split, then selectively restores the most sensitive circuits to FP16. Three findings eme
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Verktyg
- Forskning
Related AI news
- China's DeepSeek taps CITIC Securities for domestic IPOEconomic Times Tech · September 9, 2026
- Sources: China Securities Regulatory Commission is informally tightening IPO approvals for humanoid startups after a volatile debut by industry leader Unitree (The Information)Techmeme · September 9, 2026
- LG Innotek tackles glass substrate microcrack issue as 2028 production race heats upDIGITIMES · September 9, 2026
- Apple yields to memory suppliers in historic strategy shift, with Kioxia tipped as NAND long-term agreement recipientDIGITIMES · September 9, 2026
- Planning and Scheduling Business Processes under Control-Flow UncertaintyarXiv cs.AI · September 9, 2026
- 6 av 10 anställda saknar tiden innan AI fannsComputer Sweden · September 9, 2026