Characterizing the Scalability and Performance of Large-Scale AI Training Under Multi-Tenancy
Jacopo Raffi, Thomas Pasquali, Lorenzo Piarulli, Filippo Spiga, Marco Faltelli, Andreas Herten, Domenico Siracusa, Daniele De Sensi, Flavio Vella
Abstract
Characterising AI workload performance on modern HPC systems requires understanding both their scalability in isolation and their behaviour under concurrent execution. However, the interplay among parallelisation strategies, network congestion, compute capability, and interconnect technologies remains poorly understood. This work investigates the performance and scalability of AI models up to 2400 GPUs. We quantify the communication overheads and their impact across different interconnects by evaluating scale-up, scale-out, and rack-scale configurations under multiple allocation schemes. Finally, we study how multiple concurrent training jobs interfere with each other by designing a realistic noise model. We design a benchmark suite of AI models to evaluate the performance of five distinct parallelisation strategies across different supercomputing clusters, including Alps, Leonardo, LUMI, JUPITER, NVL72 GB300, and DGX A100. Our work provides a systematic characterization of the scalability and execution efficiency of distributed AI training, while offering key insights into performance behavior under realistic multi-tenant scenarios.
Create a lesson
Related papers
AceSpec: An Asymmetric Edge-Cloud Collaborative Framework for Communication-Efficient LLM Inference
Yida Zhang, Zhiyong Gao, Shuaibing Yue et al.
Federated Learning on the American Science Cloud using APPFL
Zilinghan Li, Abhijit Chunduru, Harinarayan Krishnan et al.
Towards Global Federated Genome-Wide Association Meta-Analysis Using GA4GH TES
Abhijit Chunduru, Matthew Joel, Zilinghan Li et al.
MeanField Surrogate Modeling for Scalable Runtime Scheduling of Concurrent Heterogeneous AI Inference on Shared GPUs
Youssef Ennouri, Soonhoi Ha
RT-HiSS: Ray Tracing Accelerated High Dimensional Vector Similarity Searches
Revanth Reddy Munugala, Michael Gowanlock
CREDIT: Cost-guided Reduction-reuse with Efficient DSMEM Inter-CTA Tiling
Zhengxiong Li, Tsung-Wei Huang, Umit Ogras