Work in Computational Affective Science and Computational Social Science explores a wide variety of research questions about people, emotions, behavior, and health. Such work often relies on language data that is first labeled with relevant information, such as the use of emotion words or the age of the speaker. Although many resources and algorithms exist to enable this type of labeling, discovering, accessing, and using them remains a substantial impediment, particularly for practitioners outside of computer science. Here, we present the ABCDE dataset (Affect, Body, Cognition, Demographics, and Emotion), a large-scale collection of over 400 million text utterances drawn from social media, blogs, books, and AI-generated sources. The dataset is annotated with a wide range of features relevant to computational affective and social science. ABCDE facilitates interdisciplinary research across numerous fields, including affective science, cognitive science, the digital humanities, sociology, political science, and computational linguistics.
翻译:计算情感科学与计算社会科学领域的研究探索了关于人类情感、行为与健康的多类研究问题。此类研究常需对语言数据进行标注,例如识别情感词的使用或说话者的年龄。尽管已有大量资源和算法支持此类标注,但发现、获取和运用这些工具仍构成巨大障碍,尤其对计算机科学领域外的研究者而言。本文提出ABCDE数据集(情感、身体、认知、人口统计学与情绪),这是一个包含超4亿条文本语句的大规模集合,数据源自社交媒体、博客、书籍及人工智能生成内容。该数据集标注了与计算情感科学和社会科学相关的广泛特征,为情感科学、认知科学、数字人文学、社会学、政治学及计算语言学等多个学科的交叉研究提供支撑。