Understanding as an Explicit and Assessable Component of Frontier AI Safety Decisions
Stephen Barrett, Robin Bloomfield, Alexandra Chirilă, Mamoon Masud, David Meredith Hardy, Phillip Mulvana
Abstract
Decision makers need sufficient understanding to make good decisions about training or deploying frontier AI systems. However, such decisions are increasingly made under time-pressure, and this combined with the use of AI generated artefact creation, can mean that the existence of safety cases and system cards may no longer demonstrate that sufficient understanding exists. Our provisional methodology for making understanding explicit and assessable requires the production of an explicit description of 4 objects of understanding (decision, decision-frame, safety justification, system-in-context) and a justification for the adequacy of this understanding. In addition, the methodology provides a mechanism for describing and evaluating the adequacy of the decision-maker representation of this understanding. It builds on recent developments in safety cases using the Assurance 2.0 framework to operationalise the philosophical basis of understanding from Elgin and Arendt. To assess the methodology we trialled two different scenarios. One scenario, which we investigated through role-based analysis, concerned the risk of scheming in the deployment of an AI coding agent in a robotics company and the other scenario was for the higher uncertainty, more decision-critical argument of 'If Anyone Builds It, Everyone Dies' (Yudkowsky and Soares). The trial's central finding, for these two scenarios, is that the methodology could be applied and was found to be generative: we found the analyses that justify sufficiency of understanding (internal coherence, tethering, felicitous falsehoods, external coherence) drives the engineering.
Create a lesson
Related papers
Prepared Or Unprepared? Evaluating Healthcare Workforce Readiness for Clinical Adoption of Artificial Intelligence in Nigeria
Abbas M. Rabiu, Abdulrazaq A. Zubair, Um-mulkhairi Ibrahim et al.
Could Underwater Data Centers Pose a Risk to AI Treaty Verification?
James Teague, Ashmita Rajmohan, Yannick Muehlhaeuser
Control-Theoretic Content Moderation
Benedetta Tessa, Serena Tardelli, Marco Avvenuti et al.
"If I Had to Buy Just ONE: Galaxy S26 Ultra": Auditing AI-Generated Product Recommendations
Lucas G. Uberti-Bona Marin, Thales Bertaglia, Giovanni Astante et al.
Understanding AI Provider Recommendations in Local Service Markets
Hazem Ibrahim, Yasir Zaki
Toward a Time-Aware Assessment Framework for the Carbon Cost of AI-Enabled Decarbonization
Chenrui Xu, Burcu Akinci, Christopher McComb