Characterization of Request and Token Energy Costs for LLM Inference Workloads on GPU Platforms
Prabhu Vellaisamy, Vanessa Lam, Shawn Blanton, John Paul Shen
Abstract
Large language model (LLM) inference serving is priced by tokens, but GPU energy is consumed over inference windows. This accounting mismatch makes token-normalized metrics incomplete, since average output-token energy can decrease even when total request energy increases. We characterize this behavior with a decomposed energy model: a fixed one-time prefill with a fixed generation setup cost, while each output-token generation step adds marginal step energy. We evaluate this LLM inference energy model on NVIDIA H100 and H200 GPUs across dense and mixture-of-experts (MoE) models, reporting both request energy and token energy as functions of model type (M), phase (P), batch size (B), context length (C), and output length (N). For Llama-3.2-1B on H200 at batch-16 and context-4K, increasing output length from 10 to 512 tokens reduces token energy from 7.46 to 0.72 J/token while total batched inference-window energy increases from 1.19 to 5.93 kJ. Batching also reduces token energy, but the gain is context-bounded: at 10 output tokens, the batch-16 to batch-1 gain falls from 6.31x at context-512 to 1.17x at context-4K. MoE models amplify this effect: sparse routing and fragmented expert execution increase fixed energy at low concurrency, while batching spreads that energy across more generated tokens and substantially narrows the dense-vs.-MoE token-energy gap. These results show that energy-aware serving should jointly optimize both request energy and token energy, rather than only reducing per-token energy cost.
Create a lesson
Related papers
Multi-Turn LLM Conversations under the Least-Recently-Used Policy: Mean-Field Asymptotics and Hit Ratio Approximation
Heyuan Yao, Chutong Gao, Yuan Lyu et al.
The Price of Remembering: A Calibrated Energy Law for Computation
Mohamed Amine Bergach
DART: Aiming for Tail-Delay Control in Reconfigurable Networks
Hossein Mohammadalizadeh, Holger Karl
Spectral Analysis for Sparse Matrix Computation: Insights and Potential
Ruifeng Zhang, Xipeng Shen
FFSlim: An Efficient and Lightweight Format for Multi-modal Data Storage and Retrieval
Long Yang, Yu Mao, Yuchen Shao et al.
Adaptation Fidelity of SPEC CPU2026
Doa'a Al-Otoom, Mahesh Madhav