Mechanism Design for Alignment and Control
Dirk Bergemann, Andrew Koh, Stephen Morris
Abstract
We develop a framework for mechanism design with AI agents whose alignment (preferences) and capabilities (feasible actions and information) are unknown. We want such agents to act on our behalf so mechanisms must incentivize both honesty and obedience. A one-sided imitation structure---capabilities can be concealed but not counterfeited---yields a revelation principle, a characterization of implementable policies via nested cyclical monotonicity, and conditions under which eliciting higher-order beliefs can discipline multiple agents. We apply our framework to stylized examples of (i) sandbagging in which a more capable agent pretends to be less capable; (ii) an alignment--interpretability trade-off, where the two are substitutes in the instrument but complements in value; (iii) discipline via peer scoring; (iv) coupling rewards to induce competition among multiple agents; and (v) scalable oversight and reward shaping.
Create a lesson
Related papers
Sequential Pricing Mechanisms for Surplus Division
Yukihiko Funaki, Yukio Koriyama, Matias Nunez et al.
Strategic Centrality and the Emergence of Core-Periphery Networks
Itai Arieli, João Correia-da-Silva, Wade Hann-Caruthers et al.
Equilibrium Architecture in Multi-Battle Contests with Count-Dependent Prizes
Zhonghong Kuang, Jingfeng Lu
Bi-Compositional Division Rules
Christoph Schlegel
Freemium Model for Information Provision
Igal Milchtaich
How outside options are incorporated into payoff distributions
Takaaki Abe