Verbal deception has been studied in psychology, forensics, and computational linguistics for a variety of reasons, like understanding behaviour patterns, identifying false testimonies, and detecting deception in online communication. Varying motivations across research fields lead to differences in the domain choices to study and in the conceptualization of deception, making it hard to compare models and build robust deception detection systems for a given language. With this paper, we improve this situation by surveying available English deception datasets which include domains like social media reviews, court testimonials, opinion statements on specific topics, and deceptive dialogues from online strategy games. We consolidate these datasets into a single unified corpus. Based on this resource, we conduct a correlation analysis of linguistic cues of deception across datasets to understand the differences and perform cross-corpus modeling experiments which show that a cross-domain generalization is challenging to achieve. The unified deception corpus (UNIDECOR) can be obtained from https://www.ims.uni-stuttgart.de/data/unidecor.
翻译:言语欺骗已在心理学、法医学和计算语言学领域因多种原因被研究,例如理解行为模式、识别虚假证词以及检测在线交流中的欺骗行为。不同研究领域的动机差异导致了对所研究领域的选择以及欺骗概念化上的不同,这使得难以比较模型并针对特定语言构建稳健的欺骗检测系统。通过本文,我们通过调查现有的英语欺骗数据集(包括社交媒体评论、法庭证词、特定主题的意见陈述以及来自在线策略游戏的欺骗性对话等领域的语料)来改善这一状况。我们将这些数据集整合成一个统一的语料库。基于这一资源,我们对不同数据集中欺骗的语言线索进行了相关性分析,以理解其差异,并进行了跨语料库建模实验,结果表明实现跨领域泛化具有挑战性。统一的欺骗语料库(UNIDECOR)可从 https://www.ims.uni-stuttgart.de/data/unidecor 获取。