Misinformation on YouTube is a significant concern, necessitating robust detection strategies. In this paper, we introduce a novel methodology for video classification, focusing on the veracity of the content. We convert the conventional video classification task into a text classification task by leveraging the textual content derived from the video transcripts. We employ advanced machine learning techniques like transfer learning to solve the classification challenge. Our approach incorporates two forms of transfer learning: (a) fine-tuning base transformer models such as BERT, RoBERTa, and ELECTRA, and (b) few-shot learning using sentence-transformers MPNet and RoBERTa-large. We apply the trained models to three datasets: (a) YouTube Vaccine-misinformation related videos, (b) YouTube Pseudoscience videos, and (c) Fake-News dataset (a collection of articles). Including the Fake-News dataset extended the evaluation of our approach beyond YouTube videos. Using these datasets, we evaluated the models distinguishing valid information from misinformation. The fine-tuned models yielded Matthews Correlation Coefficient>0.81, accuracy>0.90, and F1 score>0.90 in two of three datasets. Interestingly, the few-shot models outperformed the fine-tuned ones by 20% in both Accuracy and F1 score for the YouTube Pseudoscience dataset, highlighting the potential utility of this approach -- especially in the context of limited training data.
翻译:YouTube上的虚假信息是一个重要问题,亟需稳健的检测策略。本文提出了一种新颖的视频分类方法,专注于内容的真实性。通过利用视频转录本中的文本内容,我们将传统的视频分类任务转化为文本分类任务。我们采用迁移学习等先进机器学习技术来解决分类挑战。我们的方法包含两种迁移学习形式:(a)微调基础Transformer模型,如BERT、RoBERTa和ELECTRA;(b)使用句子变换器MPNet和RoBERTa-large进行少样本学习。我们将训练好的模型应用于三个数据集:(a)YouTube疫苗虚假信息相关视频,(b)YouTube伪科学视频,以及(c)假新闻数据集(文章集合)。假新闻数据集的引入将我们方法的评估范围扩展至YouTube视频之外。利用这些数据集,我们评估了模型区分有效信息与虚假信息的能力。在三个数据集中的两个上,微调模型的马修斯相关系数>0.81,准确率>0.90,F1分数>0.90。有趣的是,在YouTube伪科学数据集中,少样本模型在准确率和F1分数上均比微调模型高出20%,凸显了该方法在训练数据有限的情况下的潜在应用价值。