AI agents built on large language models can assist not only legitimate tasks but also relational manipulation. AI agents can be used to help a user maintain a deceptive identity, intensify emotional dependency, isolate a target, or prepare for later extraction. We conceptualise this risk as agentic relationship harm: workflow-level assistance that can exploit recipient vulnerability, persuasive influence, and relational power asymmetry. Existing safety evaluations and generic guardrails often treat harmfulness as a property of isolated outputs, missing role-sensitive interaction patterns. To study this, we introduce a 110-prompt benchmark with balanced attacker- and victim-side cases, a relationship-specific labelling framework, and a lightweight post-generation policy gate for local agent deployments. In our evaluation, the relationship-specific gate outperforms generic safety prompting under automated judging, with no judge-identified harmful-compliance cases on the main benchmark or multi-turn stress test while preserving victim-side protective intervention. These results suggest that relationship harm is a distinct sociotechnical risk surface and that role-sensitive evaluation plus lightweight policy gating offers a practical path beyond generic refusal prompting.
翻译:基于大语言模型的AI代理不仅可能协助合法任务,也可能助长关系操纵行为。AI代理可用于帮助用户维持欺骗性身份、强化情感依赖、隔离目标对象或为后续剥削做准备。我们将这种风险概念化为"代理关系伤害":能够在工作流层面利用接收方脆弱性、说服性影响力和关系权力不对称的助益行为。现有安全评估和通用防护措施通常将有害性视为孤立输出的属性,忽视了角色敏感的交互模式。为研究该问题,我们提出了包含110个提示的基准测试(平衡攻击方与受害方案例)、关系特定标注框架,以及适用于本地代理部署的轻量级后生成策略门控。实验表明,在自动化评估条件下,关系特定门控优于通用安全提示——主基准测试与多轮压力测试中均未出现评估者认定的有害顺从案例,同时保留了针对受害方的保护性干预措施。这些结果表明,关系伤害属于独特的社会技术风险面,而角色敏感评估结合轻量级策略门控为突破通用拒绝提示的局限提供了切实可行的路径。