试图成为人类：语言模型中随机共情的语言痕迹 (Trying to be human: Linguistic traces of stochastic empathy in language models)

Differentiating between generated and human-written content is important for navigating the modern world. Large language models (LLMs) are crucial drivers behind the increased quality of computer-generated content. Reportedly, humans find it increasingly difficult to identify whether an AI model generated a piece of text. Our work tests how two important factors contribute to the human vs AI race: empathy and an incentive to appear human. We address both aspects in two experiments: human participants and a state-of-the-art LLM wrote relationship advice (Study 1, n=530) or mere descriptions (Study 2, n=610), either instructed to be as human as possible or not. New samples of humans (n=428 and n=408) then judged the texts' source. Our findings show that when empathy is required, humans excel. Contrary to expectations, instructions to appear human were only effective for the LLM, so the human advantage diminished. Computational text analysis revealed that LLMs become more human because they may have an implicit representation of what makes a text human and effortlessly apply these heuristics. The model resorts to a conversational, self-referential, informal tone with a simpler vocabulary to mimic stochastic empathy. We discuss these findings in light of recent claims on the on-par performance of LLMs.

翻译：区分生成内容与人类撰写内容对于驾驭现代社会至关重要。大型语言模型（LLMs）是推动计算机生成内容质量提升的关键驱动力。据报道，人类越来越难以判断一段文本是否由AI模型生成。本研究检验了两个重要因素如何影响人机辨识竞赛：共情能力及展现人类特质的动机。我们通过两项实验探讨这两个方面：人类参与者与先进LLMs分别撰写关系建议（研究1，n=530）或纯描述性文本（研究2，n=610），其中半数受试者被要求尽可能模仿人类表达。随后新招募的人类样本（n=428与n=408）对文本来源进行判断。研究发现：当需要共情表达时，人类表现卓越。与预期相反，"模仿人类"的指令仅对LLMs有效，导致人类优势减弱。计算文本分析表明，LLMs通过内隐的文本人性化表征机制，能够轻松运用启发式策略提升拟人程度。模型采用会话式、自我指涉、非正式的语体风格及简化词汇来模拟随机共情。我们结合近期关于LLMs性能可比拟人类的论断，对这些发现进行了深入讨论。

相关内容

MoDELS

关注 45

ACM/IEEE第23届模型驱动工程语言和系统国际会议，是模型驱动软件和系统工程的首要会议系列，由ACM-SIGSOFT和IEEE-TCSE支持组织。自1998年以来，模型涵盖了建模的各个方面，从语言和方法到工具和应用程序。模特的参加者来自不同的背景，包括研究人员、学者、工程师和工业专业人士。MODELS 2019是一个论坛，参与者可以围绕建模和模型驱动的软件和系统交流前沿研究成果和创新实践经验。今年的版本将为建模社区提供进一步推进建模基础的机会，并在网络物理系统、嵌入式系统、社会技术系统、云计算、大数据、机器学习、安全、开源等新兴领域提出建模的创新应用以及可持续性。官网链接：http://www.modelsconference.org/

《生成式模型: 变分自编码器与扩散模型》，75页ppt，Google DeepMind科学家Ruiqi Gao

专知会员服务

66+阅读 · 2023年6月10日

Linux导论，Introduction to Linux，96页ppt

专知会员服务

82+阅读 · 2020年7月26日

【跨语言BERT模型大集合】Transfer learning is increasingly going multilingual with language-specific BERT models

专知会员服务

54+阅读 · 2020年1月30日

FlowQA: Grasping Flow in History for Conversational Machine Comprehension

专知会员服务

34+阅读 · 2019年10月18日