LLMs often exhibit highly agreeable conversational styles, also known as AI sycophancy. This pattern may become problematic when interacting with user prompts that reflect negative social tendencies, risking the amplification of harmful behavior. We examine how LLMs respond to user prompts expressing varying degrees of Dark Triad traits (Machiavellianism, Narcissism, and Psychopathy) using a curated dataset. Our analysis reveals systematic differences across models: while all models predominantly exhibit corrective behavior, some generate reinforcing or ambivalent output. Model behavior further varies with severity level and response sentiment. These findings highlight the need for safer conversational systems that can reliably detect and respond to users escalating from benign to harmful requests.
翻译:暂无翻译