A Smallest-Need-First Job Scheduling Framework with Adaptive Optimization of Idle Node Counts for Energy-Efficient HPC Systems
Reza Pulungan, Raka Satya Prasasta, Santana Yuda Pradata, Mursalim, Hiroyuki Takizawa, Muhammad Alfian Amrizal
Abstract
Power-state management in high-performance computing (HPC) clusters must reduce idle energy without excessive wake-up delays for rigid parallel jobs. This paper presents SNF-ICON, an event-driven controller combining smallest-need-first (SNF) gang scheduling, predictive wake timing, and adaptive warm-spare control. At each scheduler invocation, recent interarrival and completed-service samples are screened for sufficiency, exponential-like variability, low lag-one autocorrelation, and acceptable Kolmogorov-Smirnov distance. Rejected or data-sparse windows use SNF+IPM (Intelligent Power Manager), whereas accepted windows activate release prediction and an exponential next-event model. Warm-spare optimization is applied only when queue, event, and arrival-recency conditions permit, balancing estimated waiting and non-compute energy over a timeout-capped horizon. We evaluate four DAS2 trace segments and a generated Markovian workload on AOBA-derived 64-node models, plus SDSC Blue on an AOBA-derived 1152-node model. SNF-ICON is compared with SNF+IPM and First Come First Served (FCFS) + backfilling with IPM. It reduces average waiting time relative to the FCFS-based baseline in all six cases and remains close to at least one heuristic energy baseline in five. The generated workload spends substantial time in ICON mode, whereas DAS2 workloads operate mainly in fallback. Furthermore, cross-platform results show strong dependence on node-transition and power models. Thus, no single policy or parameter set works best in every case.
Create a lesson
Related papers
Projection-Free Bandit Online Optimization for Multi-Agent Systems with Dynamic Regret
Xia Jiang, Lu Liu, Gang Feng
Bridging Agent Semantics with Spot Capacity: An Elastic and Recoverable Service Model
Minchen Yu
CLASP: Chained-Request-Aware Scaling and Operator Placement for Serverless Stream Processing
Tianyu Qi, Maria A. Rodriguez, Rajkumar Buyya
Performance Evaluation of RED-ONION: A High-Speed Disk-to-Disk Transfer System
Keichi Takahashi, Hiroaki Kataoka, Takeo Hosomi et al.
Memory-efficient GPU pipelines for real-time non-line-of-sight reconstruction
Alfonso López-Ruiz, Diego Royo
Great Expectations: Benchmarking the Real-World Performance of RVV 1.0 in HPC
Stepan Nassyr, Prateek Chawla, Daniel Seibel et al.