Hateful comments are prevalent on social media platforms. Although tools for automatically detecting, flagging, and blocking such false, offensive, and harmful content online have lately matured, such reactive and brute force methods alone provide short-term and superficial remedies while the perpetrators persist. With the public availability of large language models which can generate articulate synthetic and engaging content at scale, there are concerns about the rapid growth of dissemination of such malicious content on the web. There is now a need to focus on deeper, long-term solutions that involve engaging with the human perpetrator behind the source of the content to change their viewpoint or at least bring down the rhetoric using persuasive means. To do that, we propose defining and experimenting with controllable strategies for generating counter-arguments to hateful comments in online conversations. We experiment with controlling response generation using features based on (i) argument structure and reasoning-based Walton argument schemes, (ii) counter-argument speech acts, and (iii) human characteristics-based qualities such as Big-5 personality traits and human values. Using automatic and human evaluations, we determine the best combination of features that generate fluent, argumentative, and logically sound arguments for countering hate. We further share the developed computational models for automatically annotating text with such features, and a silver-standard annotated version of an existing hate speech dialog corpora.
翻译:仇恨言论在社交媒体平台上普遍存在。尽管自动检测、标记和屏蔽此类虚假、冒犯性和有害网络内容的工具近年已趋于成熟,但仅靠这种被动且强制的方法只能提供短期和表面的补救措施,而施害者依然持续作恶。随着能够大规模生成流畅合成且引人入胜内容的大型语言模型公开可用,人们担忧这类恶意内容在网络上的传播速度将急剧增长。当前需要关注更深层次的长期解决方案,这些方案需与内容源头背后的人类施害者进行互动,通过说服性手段改变其观点,或至少降低其言论的攻击性。为此,我们提出定义并实验可控策略,用于在在线会话中生成针对仇恨言论的反驳论点。我们基于以下特征实验性地控制回复生成:(i)基于论证结构与推理的沃尔顿论证模式,(ii)反驳性言语行为,以及(iii)基于人类特质(如大五人格和人类价值观)的个性特征。通过自动评估和人工评估,我们确定了生成流畅、具有论证性且逻辑合理的反制仇恨论点时最优的特征组合。我们进一步分享了用于自动标注文本中此类特征的计算模型,以及一个现有仇恨言论对话语料库的银标准标注版本。