Recent developments in generative AI have shone a spotlight on high-performance synthetic text generation technologies. The now wide availability and ease of use of such models highlights the urgent need to provide equally powerful technologies capable of identifying synthetic text. With this in mind, we draw inspiration from psychological studies which suggest that people can be driven by emotion and encode emotion in the text they compose. We hypothesize that pretrained language models (PLMs) have an affective deficit because they lack such an emotional driver when generating text and consequently may generate synthetic text which has affective incoherence i.e. lacking the kind of emotional coherence present in human-authored text. We subsequently develop an emotionally aware detector by fine-tuning a PLM on emotion. Experiment results indicate that our emotionally-aware detector achieves improvements across a range of synthetic text generators, various sized models, datasets, and domains. Finally, we compare our emotionally-aware synthetic text detector to ChatGPT in the task of identification of its own output and show substantial gains, reinforcing the potential of emotion as a signal to identify synthetic text. Code, models, and datasets are available at https: //github.com/alanagiasi/emoPLMsynth
翻译:生成式AI的最新发展使高性能合成文本生成技术备受关注。这类模型如今广泛可用且易于使用,凸显了对同等强大合成文本识别技术的迫切需求。基于此,我们从心理学研究中汲取灵感——研究表明人类可能受情感驱动并在其撰写的文本中编码情感。我们假设预训练语言模型存在情感缺陷,因为它们在生成文本时缺乏这种情感驱动力,因此可能生成具有情感不连贯性(即缺乏人类撰写文本中存在的情感连贯性)的合成文本。随后,我们通过在情感任务上微调预训练语言模型,开发了一种情感感知检测器。实验结果表明,我们的情感感知检测器在多种合成文本生成器、不同规模模型、数据集和领域上均取得了性能提升。最后,我们将情感感知合成文本检测器与ChatGPT在识别自身输出任务中进行比较,展示了显著优势,进一步强化了情感作为合成文本识别信号的潜力。代码、模型和数据集可在https://github.com/alanagiasi/emoPLMsynth获取。