The rapid advancement of large language models (LLMs) has opened new possibilities for AI for good applications. As LLMs increasingly mediate online communication, their potential to foster empathy and constructive dialogue becomes an important frontier for responsible AI research. This work explores whether LLMs can serve not only as moderators that detect harmful content, but as mediators capable of understanding and de-escalating online conflicts. Our framework decomposes mediation into two subtasks: judgment, where an LLM evaluates the fairness and emotional dynamics of a conversation, and steering, where it generates empathetic, de-escalatory messages to guide participants toward resolution. To assess mediation quality, we construct a large Reddit-based dataset and propose a multi-stage evaluation pipeline combining principle-based scoring, user simulation, and human comparison. Experiments show that API-based models outperform open-source counterparts in both reasoning and intervention alignment when doing mediation. Our findings highlight both the promise and limitations of current LLMs as emerging agents for online social mediation.
翻译:大语言模型的快速发展为人工智能向善应用开辟了新的可能性。随着大语言模型日益介入在线交流,其促进共情与建设性对话的潜力已成为负责任人工智能研究的重要前沿。本研究探讨大语言模型能否不仅作为检测有害内容的审核者,更能成为理解并化解在线冲突的调解者。我们的框架将调解分解为两个子任务:判断(大语言模型评估对话的公平性与情感动态)与引导(模型生成具有共情性、能缓和冲突的信息以引导参与者走向和解)。为评估调解质量,我们构建了一个基于Reddit的大规模数据集,并提出一个结合基于原则的评分、用户模拟和人工对比的多阶段评估流程。实验表明,在执行调解任务时,基于API的模型在推理与干预对齐方面均优于开源模型。我们的研究结果既揭示了当前大语言模型作为在线社会调解新兴代理的潜力,也指出了其局限性。