OverAct: Measuring and Mitigating Proactive Over-Authorization in LLM Tool-Calling Agents
Taolin Zhang, Jiuheng Wan, Hanyu Wang, Tingyuan Hu, Chengyu Wang
Abstract
LLM agents with tool-calling capabilities can access external services and private user data, but they may retrieve more information than a user's request explicitly requires. We study this behavior in structured tool-calling agents and term it proactive over-authorization. This setting differs from filesystem-level coding agents because the main risk is unnecessary access to private data. We introduce OverAct, a controlled benchmark spanning eight privacy-sensitive domains with deterministic, judge-free scoring, together with an interpretive decision-theoretic framework that yields three testable predictions. Across seven models from four families, all models significantly exceed authorized scope. Request specificity is the strongest predictor of severity, over-authorization grows sublinearly with tool-pool size, and decoding temperature has little effect. These patterns are consistent with a cost-asymmetry account, suggesting that over-authorization arises more from structural decision tendencies than from decoding randomness. We also propose SelfAudit, a zero-shot inference-time method that generates request-grounded justifications and filters unjustified calls before execution. Ablation shows that explicit filtering is the main driver of scope reduction. SelfAudit reduces privacy-oriented excess by 43% without oracle knowledge.
Create a lesson
Related papers
System-Level Optimization Beyond Cryptographic Kernels: An ML-KEM Case Study on Arm Cortex-M7
Mahmoud Abdelhafeez Sayed, Mostafa Taha, Gurp Nijjer
A Hybrid Approach to Malware Detection: Integrating Few-Shot Model-Agnostic Meta-Learning with Autoencoders
Emmanuela Andam, Yasir Abbas Zaidi, Abdelali Hadir et al.
Detection and Resolution of Periodic Artifacts in OpenDP's Discrete Laplace Sampler
Cesare Gerolimetto Fabrello, Valeria Rossi, Alberto Trombetta et al.
A Structured State Space Sequence Model for Multi-Class Classification of Malware
Emmanuela Andam, Rana Shaaban, Emanuel Grant et al.
From Network Intrusion Detection to Blockchain-Backed Endpoint Detection and Response: Mapping the Landscape of Decentralized Detection-and-Response Architectures
Yahya Shahsavari, Sara Rouhani, Kaiwen Zhang
Walking the Embedding Space: Datastore Extraction from Multimodal RAG
Maria Carmen Jica, Ali Satvaty, Suzan Verberne et al.