Application Failures and Machine Computational Efficiency
Carlo Graziani, Bethany Lusch, O. E. Bronson Messer
Abstract
We present a framework for evaluating uptime efficiency of Exascale-class scientific computers when application failure rates are appreciable. This is the situation that confronts current leadership-class scientific computing platforms and large AI training installations. What distinguishes scientific computing platforms is the heterogeneity of their applications. We argue that this diversity requires that failure rates and mean intervals between failures should be specified in terms of usage (e.g. node-hours) rather than time, as is currently customary. We consider the usage loss terms due to failures, to checkpointing, and to restart costs, and update the framework of Daly (2006) allowing users to specify optimal checkpointing usage intervals that minimize such losses. We derive the machine computational efficiency, which specifies the expected fractional resource allocation that is available for scientific computation. We illustrate the methodology using one year of production runtime data from the Frontier supercomputer at Oak Ridge National Laboratory.
Create a lesson
Related papers
Replication-Aware Placement of Functions and Data in the Edge-Cloud Continuum
Dario d'Abate, Matteo Cenzato, Matteo Briscini et al.
Ermes: a Stateful Serverless Platform for the Edge-to-Cloud Continuum
Matteo Cenzato, Dario d'Abate, Arianna Dragoni et al.
Fluid Notarization: Verifiable Evolution of Concurrently Edited Structured Documents
Amos Brocco, Giuliano Gremlich, Roberto Guidi
Ask the Tool, Don't Guess: Agent Tool Calls Hold Their Progress, and the Serving System Should Read It
Yipeng Liu, Yingqiang Zhang, Feifei Li et al.
A Distributed Computing Framework for Satellite Swarms
Ezra Fielding, Clement Demazure, Guthemberg Silvestre et al.
Vigil: Accountable Liveness against Selective Silence
Jiawei Cheng, Huiping Sun, Rui Zhou et al.