Request-Level Energy Attribution for Batched LLM Serving
arXiv cs.AIen
arXiv:2608.00026v1 Announce Type: new Abstract: Batched LLM serving improves throughput but complicates energy accounting. GPU power telemetry is aggregate, whereas sustainability reporting, chargeback, and workload analysis often require request-level energy charges. Existing inference-energy benchmarks report model-, phase-, or token-level energy, and recent carbon-accounting work motivates Shapley fairness conceptually. Neither provides measured request-level ground truth, so how far the accounting rules used in practice deviate from a fair allocation has remained unknown. We present JouleShare, an attribution framework with two components. An offline harness establishes this ground truth
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Verktyg
- Forskning
Related AI news
- Högre chefer missbrukar skugg-AI dubbelt så ofta som vanliga anställdaComputer Sweden · August 4, 2026
- Avec l’introduction des « aperçus IA », « le Web et les applications mobiles pourraient n’avoir été qu’une étape intermédiaire dans la transformation numérique »Le Monde Pixels · August 4, 2026
- Enhancing LLMs with Context-Specific Knowledge for Mitigating Misinformation in SMEs: A RAG-based Modeling and AnalysisarXiv cs.AI · August 4, 2026
- Linguistic Context Recodes Visual Representations in Vision-Language ModelsarXiv cs.AI · August 4, 2026
- Can LLM Agents Price Competitively? A Dynamic Multi-Attribute Auction Benchmark for Agentic CommercearXiv cs.AI · August 4, 2026
- CrystalMem: Elastic Memory for Self-Evolving LLM Agents via Knowledge CrystallizationarXiv cs.AI · August 4, 2026