Capability-Gated Language Models: Security Composes, Utility Does Not
Patrikas Vanagas, Augustas Mačijauskas, Laurynas Lopata
Abstract
Deployed language model safeguards (safety fine-tuning, filtering, unlearning) vary by principal only outside the model weights: filters are reconfigured, tiers are multiplied, and artefacts are reissued; inside one set of weights every request meets the same model configuration. This motivates us to define capability-gated deployment: per-principal access control inside one set of weights, whose configurations form a lattice - meets accumulate a principal's restrictions and joins pool a coalition's reach. We instantiate it by sparse rank gating over an existing nested-factorisation mechanism, guide profile search with one-pass attribution, and read every result once from a pre-registered held-out split. Security composes: provably at meets under a monotone-elicitation assumption we falsify pointwise. In two lineages the median held-out meet deepens suppression; the one effect surviving correction strengthens it. Utility does not: individually harmless profiles can compose to retention and fluency damage, and no compositional bound exists.
Create a lesson
Related papers
Don't Trust the Code, Check Its Effects: Runtime Refinement for Regenerated Systems Code Under an Adversarial Generator
Jinhao Hu, Ashvin Goel, Laurent Bindschaedler
NeuroPriv: Adversarial Representation Learning for Privacy in Wearable EEG Systems
Sarmistha Sarna Gomasta, Bhawana Chhaglani, Prashant Shenoy
OreProof: Verifiable Provenance with Limited Disclosure for Critical-Minerals Supply Chains Using Zero-Knowledge Proofs
Oleksandr Hrabar, Hossein Arshadi Soufiani, Henry M. Kim et al.
Workload Identification with Physical Side Channels for AI Governance
Simone Gargiulo, Gabriel Kulp
Delegation Without Trust: An Empirical Gap Analysis of Identity, Authorization, and Runtime Governance in Multi-Agent LLM Systems
Panduranga Sai Varma Dantuluri, Jyotirmoy Sundi
DUPIN: Attack Learning Is Still Needed! Demonstrating Few-Shot after Unsupervised Pretraining Is A Nimble Forensics Learner
Chanwoo Bae, Hailun Ding, Shiqing Ma et al.