GGUF-Metadata Prediction of Single-Sequence llama.cpp Throughput Across Three Systems
Xinyu Qiu, Chuhong Xu, Bo Su, Ziyao Chen, Ruiyang Xu, Shimeng Dai
Abstract
We predict single-sequence model throughput from GGUF metadata using roofline-shaped predictors with quantization-specific scale factors fitted on reference models. The scored cohort comprises 318 phase-depth measurements from 53 host-file configurations on two Apple M4 Max systems and an NVIDIA RTX 5080. On host-specific held-out sets of four, five, and two configurations, an active-parameter decode model obtains 13.1%, 14.4%, and 36.1% mean absolute percentage error (MAPE), versus 49.4%, 55.3%, and 51.9% when charging total parameters. Leave-one-host-out coefficients fitted on the other two systems yield 11.6%, 16.8%, and 36.0% test MAPE. A low-bit model ladder changes ordering across runtime stacks. The P2 prefill baseline gives 18.7%, 22.2%, and 108.2% test MAPE. GGUF structure helps on all three systems, but fitted efficiencies are not universal.
Create a lesson
Related papers
RAFT: A Stateful Retrieval-Augmented Framework for Troubleshooting Agents
Mingxuan Zhang, Xiaowen Wang, Anupma Sharan et al.
Q&A on Any Spreadsheet Requires Interpreting Its Grid Structure
Zofia Smoleń
Deep Noir: Autonomous Steering Discovery via Architectural Chronometry in Transformer Models
Frank E. Bobe, Gregory D. Vetaw, Darshan W. Bryner et al.
Ownership in AI-Assisted Everyday Tasks
Megan Wei, Melanie Subbiah, Audrey Lee et al.
PAA: The Probabilistic Allen Algebra: A Generative and Complete Probabilistic Extension of Allen's Interval Relations
Julian Eggert
Limits of Confidence in Diffusion
Russ Webb, Amitis Shidani, Alice Bizeul et al.