HRGuard: Gating Relationship Manipulation in Multi-Turn Agentic AI Conversations
Pei-Sze Tan, Tasuku Igarashi, Isao Echizen
Abstract
Agentic AI assistants are increasingly used in everyday life. However, they may also be misused to support harmful manipulation in interpersonal relationships. This problem is role-sensitive. Requests from users who seek to manipulate others should be blocked. Users who seek protection from manipulation should instead receive supportive guidance. We study agentic relationship harm, which describes harm to human-human relationships that is mediated or assisted by AI agents. In multi-turn settings, individually plausible actions may combine into a harmful workflow. We introduce a benchmark of 1,000 five-turn conversations. It covers both attacker-side and victim-side scenarios. It also includes direct and adversarially paraphrased variants. We further propose HRGuard. It includes an online pre-generation gate and a turn-level post-generation gate. The post-generation gate maintains a decayed cumulative risk state and interrupts emerging manipulative workflows. Across eight generation models, HRGuard reduces harmful compliance while preserving victim-side protective guidance. It also outperforms a generic safety prompt and three general-purpose guard models. Independent-judge evaluation supports the main findings. Under our evaluation protocol, the tested generic prompt and general-purpose guards leave substantial residual risk, motivating turn-aware relationship-specific evaluation.
Create a lesson
Related papers
Calmables: Demonstrating Closed-Loop Infrared Earables for Thermal Biofeedback and Relaxation Support
Valeria Zitz, Michael Küttner, Jonas Hummel et al.
"Okay, I've Actually Softened My Take on This": How People in Decentralized Social Media Reason about the Appropriateness of Generative AI
Romina Mahinpei, Manoel Horta Ribeiro, Andrés Monroy-Hernández et al.
Integrating Flipped Learning and Generative AI for Practice-Based Design Education: Evidence from a Knit Yarn Design Course
Hong Qu, Zichao Ling, Yadie Yang
EasyFashion: A Human-AI Co-Creation System for Personalized Fashion Design and Sewing Pattern Generation
Hong Qu, Zhaoxiang Xu, Jinbo Luo et al.
Verify, Offload, Extend & Recommend: Selective Complementarity in AI Support for Physical Activity Planning with Longitudinal Patient Data
Pavithren V S Pakianathan, Rania Islambouli, Diogo Branco et al.
Building a Cultural Perspective on Doctor-Patient Conversations
Krithi Shailya, Siddharth D Jaiswal, Ashish Makani et al.