This study reviews and compares methods for single-label and multi-label text classification, categorized into bag-of-words, sequence-based, graph-based, and hierarchical methods. The comparison aggregates results from the literature over five single-label and seven multi-label datasets and complements them with new experiments. The findings reveal that all recently proposed graph-based and hierarchy-based methods fail to outperform pre-trained language models and sometimes perform worse than standard machine learning methods like a multilayer perceptron on a bag-of-words. To assess the true scientific progress in text classification, future work should thoroughly test against strong bag-of-words baselines and state-of-the-art pre-trained language models.
翻译:本研究回顾并比较了单标签和多标签文本分类方法,这些方法被分为词袋模型、基于序列、基于图以及基于层次结构的方法。比较结果整合了文献中五个单标签和七个多标签数据集上的实验结果,并通过新实验进行了补充。研究结果显示,所有近期提出的基于图和基于层次结构的方法均未能超越预训练语言模型,有时甚至不如词袋模型上的多层感知器等标准机器学习方法。为评估文本分类领域的真实科学进展,未来研究应针对强词袋基线模型和最新预训练语言模型进行充分测试。