Adaptation Fidelity of SPEC CPU2026
Doa'a Al-Otoom, Mahesh Madhav
Abstract
Standardized benchmarks are often criticized for not being "real workloads," but this critique is rarely backed by data. This paper provides the first systematic, quantitative analysis of the "fidelity gap" between the SPEC CPU2026 suite and its original, upstream open-source counterparts. We compile both the SPEC benchmarks and their upstream applications and execute them with official input workloads under two scenarios: a single-copy latency run and a 192-copy throughput run. Our findings show that most benchmarks exhibit high fidelity in single-copy runs, while a few outliers reveal the impact of SPEC's adaptation process. The multi-copy results further highlight the necessity of this adaptation: several benchmarks become significantly more efficient than their upstream versions under heavy load, underscoring the importance of I/O reduction. This work offers data-driven validation of SPEC's methodology, showing that the fidelity gap is not a flaw but a quantifiable consequence of enforcing portability, determinism, and CPU-centric measurement.
Create a lesson
Related papers
Multi-Turn LLM Conversations under the Least-Recently-Used Policy: Mean-Field Asymptotics and Hit Ratio Approximation
Heyuan Yao, Chutong Gao, Yuan Lyu et al.
The Price of Remembering: A Calibrated Energy Law for Computation
Mohamed Amine Bergach
DART: Aiming for Tail-Delay Control in Reconfigurable Networks
Hossein Mohammadalizadeh, Holger Karl
Spectral Analysis for Sparse Matrix Computation: Insights and Potential
Ruifeng Zhang, Xipeng Shen
Characterization of Request and Token Energy Costs for LLM Inference Workloads on GPU Platforms
Prabhu Vellaisamy, Vanessa Lam, Shawn Blanton et al.
FFSlim: An Efficient and Lightweight Format for Multi-modal Data Storage and Retrieval
Long Yang, Yu Mao, Yuchen Shao et al.