How to Catch a GPU: A Taxonomy of Verification and Enforcement Mechanisms for International AI Agreements
Raymond Koopmanschap, Otto Barten
Abstract
Several international agreements have been proposed to regulate frontier AI development in response to catastrophic risks. However, there is no structured way to evaluate whether these proposals are enforceable, to assess where they might fail in practice, or to determine which combination of policies is most effective. We propose a taxonomy based on the principle that wherever sufficient capacity exists to violate an agreement, it must be under a control regime. This decomposes the problem of ensuring compliance with the agreement into preventing uncontrolled resource acquisition, detecting all capacity outside the control regime, and preventing escape from the control regime. Existing proposals consist of individual policies that address one or more of these sub-problems. Because the compute required for dangerous capabilities may decrease over time, more actors can violate an agreement and enforcement of these policies becomes harder. We define the enforcement breaking point as the FLOP-threshold or equivalent metric at which a policy loses its effectiveness in solving the sub-problem, and introduce a set of factors to assess how and why this breakdown occurs. This reveals which of the three sub-problems any given proposal adequately addresses, and where enforceability breaks down first. Applying this taxonomy to existing proposals reveals that more focus is placed on preventing escape from the control regime, while preventing resource acquisition and detecting all capacity outside the control regime receive less attention. By making these gaps explicit, this taxonomy can help researchers and policymakers prioritize future enforcement and verification efforts.
Create a lesson
Paper details
20 pages, 1 figure. Accepted at International Conference on Large-Scale AI Risks 2026 at KU Leuven