MeanField Surrogate Modeling for Scalable Runtime Scheduling of Concurrent Heterogeneous AI Inference on Shared GPUs
Youssef Ennouri, Soonhoi Ha
Abstract
Deploying heterogeneous AI models concurrently on a shared GPU introduces resource contention that complicates runtime scheduling. While surrogate models avoid costly online benchmarking, their profiling requirements typically grow combinatorially with the number of co-running models, limiting scalability. We propose a MeanField surrogate that predicts per-model performance from local configuration and aggregate GPU state rather than explicitly modeling all joint interactions. Experiments on concurrent LLM and vision workloads across N ∈ \2,3,4,5,6\ show high predictive accuracy (R2 ≈ 0.96) with an empirical sample budget that grows approximately linearly in N, in contrast to the combinatorial cost of fully joint profiling. Integrated into a genetic algorithm scheduler, the surrogate scales to an N=5 problem with 78,732 feasible joint configurations, remaining within 0.10% of the exhaustive search with zero SLA violations across eight dynamic workload scenarios, while complete online GA decisions take 26 ms median, about 5× faster than exhaustive surrogate search.
Create a lesson
Related papers
AceSpec: An Asymmetric Edge-Cloud Collaborative Framework for Communication-Efficient LLM Inference
Yida Zhang, Zhiyong Gao, Shuaibing Yue et al.
Federated Learning on the American Science Cloud using APPFL
Zilinghan Li, Abhijit Chunduru, Harinarayan Krishnan et al.
Towards Global Federated Genome-Wide Association Meta-Analysis Using GA4GH TES
Abhijit Chunduru, Matthew Joel, Zilinghan Li et al.
RT-HiSS: Ray Tracing Accelerated High Dimensional Vector Similarity Searches
Revanth Reddy Munugala, Michael Gowanlock
CREDIT: Cost-guided Reduction-reuse with Efficient DSMEM Inter-CTA Tiling
Zhengxiong Li, Tsung-Wei Huang, Umit Ogras
Scaling Inference Prefill with High-Radix Photonic Interconnects
Arulselvan Madhavan, Peter Carson, Taylor Groves et al.