This paper presents a new approach to the FNC-1 fake news classification task which involves employing pre-trained encoder models from similar NLP tasks, namely sentence similarity and natural language inference, and two neural network architectures using this approach are proposed. Methods in data augmentation are explored as a means of tackling class imbalance in the dataset, employing common pre-existing methods and proposing a method for sample generation in the under-represented class using a novel sentence negation algorithm. Comparable overall performance with existing baselines is achieved, while significantly increasing accuracy on an under-represented but nonetheless important class for FNC-1.
翻译:论文摘要:本文提出了一种新的FNC-1虚假新闻分类任务方法,该方法利用来自相似自然语言处理任务(即句子相似度和自然语言推理)的预训练编码器模型,并基于此提出了两种神经网络架构。我们探索了数据增强方法,以解决数据集中的类别不平衡问题,采用现有常用方法,并提出了一种利用新颖句子否定算法生成欠表示类别样本的方法。所提方法在整体性能上与现有基线相当,同时显著提高了FNC-1中一个欠表示但重要类别的分类准确率。