Exploiting the Interplay of Compute- and Memory-Bound kernels in MPI Applications
Ayesha Afzal, Krishna Manda, Georg Hager
Abstract
Parallel applications are often designed for synchronous, lock-step execution, treating communication stalls as performance hazards. Yet, in a communication-light application without frequent synchronization points that alternates between compute-bound memory-bound execution, an MPI communication stall can act as an unintentional relief on memory-bandwidth contention. We demonstrate this using a Parallel Optical Flow Solver, which combines a compute-bound Ray Tracing kernel with a memory-bound Optical Flow Solver kernel and negligible inter-process communication. This program shows considerable speedup via desynchronization and automatic overlap between compute- and memory-bound phases, showing that natural desynchronization is an architecture-aware optimization. An optimal speedup is achieved when the number of processes concurrently executing the memory-bound phase on a ccNUMA domain is near the bandwidth saturation point. We also show a case where reducing communication overhead using MPI asynchronous progress significantly degrades performance because it allows too many ranks to contend for memory bandwidth simultaneously. In order to study the dynamics under more controlled conditions, we develop a tunable dual-kernel microbenchmark, with which we show that significant application or system noise (natural or injected) is required to achieve full desynchronization. Finally, we also validate these results using a bandwidth-aware, model-based simulator.
Create a lesson
Related papers
MoE-CORE: Coordinated Expert Offloading and Residency for Memory-Constrained MoE Inference
Ke Yang, Yongji Gao, Xushi Li et al.
ePACT: Energy-Performance-Aware Commitment Tracking for LLM Serving
You Peng, Youhe Jiang, Chen Wang et al.
Towards a Cloud Fog Edge System for Smart Building
Christophe Cérin, Mamadou Sow, Frédéric Andrès
GridSMR: Causal Compression for Sharded Blockchains
Shir Cohen, Adam Alon, Raz Omessi et al.
GPU-Initiated Communication: Dissecting Down to the Bone
Javid Baydamirli, Ismayil Ismayilov, Kaan Oktay et al.
Blockchain Lifecycle Prediction - Dead Coins
Uwe A. Kuehn, Syed Muhammad Adnan