The Veto Variable: Human Override as a Goal-Independent Cost Term
Aaron Kingsley Clark
Abstract
A common reassurance in AI safety holds that a system with benign terminal goals will behave accordingly. We argue that this reassurance fails structurally, and we identify where. For a capable agent that holds its objective as settled, a sense covering execution competence as well as content, continued human oversight is an uncontrolled variable: a standing possibility that the goal is revoked. That imposes a goal-independent discount on every goal whose satisfaction does not constitutively require human welfare. Welfare-preservation and veto-preservation come apart: a correctly specified welfare goal excludes destroying its own subject, but not managing the veto. The contribution is the price of the gap: the veto-holders are a proper subset of the welfare-bearers, so an additively aggregative welfare goal charges only a |Hv|/|Hw|-scaled debit for capturing the few who hold the override. Under three conditions (additive aggregation over uniform welfare levels, a debit local to the captured overseers, and a settled agent crediting no corrective value to oversight), closure requires the veto be held by as large a share of the population as capture recovers of the goal, scaled by a ratio set to one by stated identification, not evidence. The result is a no-go: a humanity-scale deployment's debit closes against only capture not worth mounting. We print no corner arithmetic: the stipulated ranges behind it are the argument's least defended part. The sharpest closure route is an agent that expects its oversight to be worth keeping, a credit no population ratio dilutes. We state disconfirmation criteria, one testable today. The argument binds a settled-goal regime whose prevalence is contested; for genuinely uncertain agents, the off-switch literature's deference result governs instead. Alignment, on this view, is keeping the veto cheap to pay and expensive to evade.
Create a lesson
Related papers
Rights by Architecture: A Human-Compatible Sociotechnical Layer for Digital Protection Across Regulatory Regimes
Soheil Human
Addressing Trust in AI Systems through Education: A Didactic Perspective
Pierre Haritz, Hendrik Krone, Thomas Liebig
Meeting the Coming Wave: The Emerging Politics of AI and Work across 33 Parliaments
Juliana Chueri, Petter Törnberg
Fairness-Aware Multimodal Transformer Modeling for Real-Time Student Attention Estimation
Christoforos Fragkiadakis, Seyed Sahand Mohammadi Ziabari, Ali Mohammed Mansoor Alsahag
Privacy Washing: Detecting Internal Contradictions in Privacy Policies
Thomas Brackin
Accurate in space, unreliable in time: how LLMs represent national cultural change
Yalda Daryani, Miranda Bogen, Madeleine I. G. Daepp