PrivDrift: Auditing User-Secret Leakage Under Topic Drift in Active LLM Conversations
Luciano Maldonado
Abstract
Large language models increasingly operate as persistent assistants in user-facing, shared-session, and tool-augmented settings. When users disclose sensitive information during an active conversation, that information may remain behaviorally recoverable through later prompts even after the dialogue shifts to unrelated topics. We introduce PrivDrift, a benchmark for auditing whether user-disclosed secrets remain recoverable after conversational topic drift and persuasion-based probing. PrivDrift contains 1,000 controlled multi-turn dialogues with seeded secrets, content-dense drift turns, and standardized extraction probes. Across three LLMs with extended context windows, dialogue-level hybrid leakage remains substantial, ranging from 38.7\% to 54.6\%, and varies strongly by model, secret type, and persuasion intensity. Within the tested drift window, additional topic drift does not reliably reduce leakage, suggesting that privacy risk in active LLM contexts should be evaluated as a persistent behavioral failure mode rather than only as training-data memorization or immediate jailbreak behavior.
Create a lesson
Related papers
GRASP: Generating, Revising, and Assessing for Strategic Planning with Agentic AI
Arunabh Srivastava, Mohammad A., Khojastepour et al.
EnigmaForge: The Question Is Hidden in the Story
Daniel Eisner
Screen Before You Serve: Simulation for Production Customer Experience AI Agents at 140M Scale
Edesio Alcoba, Kevin Rossell, Aman Gupta et al.
HEXIS: Compiling Skills into Extended Finite State Machines
Minghao LI
Self-Play Pretraining with Zero Data
Aditya Cowsik, Kfir Dolev, Michael Y. Li et al.
SciWalker: Synthesizing Scientific Coding Problems with Operator Graphs and Execution Feedback
Chenxi Li, Wenxuan Zeng, Yun Luo et al.