GGUF-Metadata Prediction of Single-Sequence llama.cpp Throughput Across Three Systems

| Source: arXiv AI

Tags: llama.cpp, GGUF, LLM inference, Apple Silicon, M4 Max, RTX 5080, quantization

llama.cpp inference throughput on Apple M4 Max can be predicted from GGUF file metadata with 13–14% mean error by counting active parameters rather than total parameters — 3–4x more accurate than the naive total-parameter baseline, with weaker results on RTX 5080 at 36% error.

Details

Qiu et al. tackle a practical problem for local LLM deployment: predicting how fast a model will run without actually running it. Using roofline-shaped predictors fitted on reference models, they predict single-sequence throughput from GGUF metadata — the structured model file format used by llama.cpp. The dataset comprises 318 phase-depth measurements from 53 host-file configurations on two Apple M4 Max systems and an NVIDIA RTX 5080. The key insight is that active-parameter counting outperforms total-parameter counting dramatically: on held-out sets, an active-parameter decode model achieves 13.1–14.4% MAPE on Apple hardware versus 49.4–55.3% when using total parameters. RTX 5080 results are weaker at 36% MAPE, likely due to GPU memory bandwidth dynamics not fully captured by the roofline model. Leave-one-host-out validation (fitting on two systems, testing on the third) yields similar results, confirming modest generalizability across hardware of the same type. For practitioners deploying quantized models locally, this provides a lightweight planning tool: look up model metadata and estimate throughput before downloading or running. A low-bit model ladder changes performance ordering across runtime stacks — Q4 may outperform Q8 on some hardware. Submitted to ICASSP 2027.