To analyse large numbers of texts, social science researchers are increasingly confronting the challenge of text classification. When manual labeling is not possible and researchers have to find automatized ways to classify texts, computer science provides a useful toolbox of machine-learning methods whose performance remains understudied in the social sciences. In this article, we compare the performance of the most widely used text classifiers by applying them to a typical research scenario in social science research: a relatively small labeled dataset with infrequent occurrence of categories of interest, which is a part of a large unlabeled dataset. As an example case, we look at Twitter communication regarding climate change, a topic of increasing scholarly interest in interdisciplinary social science research. Using a novel dataset including 5,750 tweets from various international organizations regarding the highly ambiguous concept of climate change, we evaluate the performance of methods in automatically classifying tweets based on whether they are about climate change or not. In this context, we highlight two main findings. First, supervised machine-learning methods perform better than state-of-the-art lexicons, in particular as class balance increases. Second, traditional machine-learning methods, such as logistic regression and random forest, perform similarly to sophisticated deep-learning methods, whilst requiring much less training time and computational resources. The results have important implications for the analysis of short texts in social science research.
翻译:为分析大量文本,社会科学研究者正日益面临文本分类的挑战。当人工标注不可行,研究者需寻求自动化文本分类方法时,计算机科学提供了机器学习方法的实用工具箱,但这些方法在社会科学领域的表现尚缺乏充分研究。本文通过将最广泛使用的文本分类器应用于社会科学研究的典型场景——一个类别出现频率较低、规模较小的标注数据集(作为大型未标注数据集的组成部分),比较了它们的性能。以Twitter上关于气候变化的交流为例,这是跨学科社会科学研究中日益受关注的议题。我们使用包含来自各国际组织的5750条推文的新数据集(涉及高度模糊的气候变化概念),评估了基于是否涉及气候变化主题自动分类推文的方法性能。在此背景下,我们强调两个主要发现:第一,监督式机器学习方法的表现优于当前最先进的词汇表方法,尤其在类平衡度提高时更为明显;第二,逻辑回归和随机森林等传统机器学习方法的表现与复杂的深度学习方法相当,但所需训练时间和计算资源显著减少。该结果对社会科学研究中的短文本分析具有重要启示意义。