Threat actors continue to exploit geopolitical and global public events launch aggressive campaigns propagating disinformation over the Internet. In this paper we extend our prior research in detecting disinformation using psycholinguistic and computational linguistic processes linked to deception and cybercrime to gain an understanding of the features impact the predictive outcome of machine learning models. In this paper we attempt to determine patterns of deception in disinformation in hybrid models trained on disinformation and scams, fake positive and negative online reviews, or fraud using the eXtreme Gradient Boosting machine learning algorithm. Four hybrid models are generated which are models trained on disinformation and fraud (DIS+EN), disinformation and scams (DIS+FB), disinformation and favorable fake reviews (DIS+POS) and disinformation and unfavorable fake reviews (DIS+NEG). The four hybrid models detected deception and disinformation with predictive accuracies ranging from 75% to 85%. The outcome of the models was evaluated with SHAP to determine the impact of the features.
翻译:威胁行为体持续利用地缘政治和全球公共事件,在互联网上发起传播虚假信息的恶意活动。本文扩展了我们先前利用与欺骗和网络犯罪相关的心理语言学及计算语言学过程来检测虚假信息的研究,旨在深入理解影响机器学习模型预测结果的特征。本文尝试通过使用eXtreme Gradient Boosting机器学习算法,在基于虚假信息与诈骗、虚假正面与负面在线评论或欺诈数据训练的混合模型中,确定虚假信息中的欺骗模式。我们构建了四个混合模型:基于虚假信息与欺诈数据训练的模型(DIS+EN)、基于虚假信息与诈骗数据训练的模型(DIS+FB)、基于虚假信息与虚假好评数据训练的模型(DIS+POS)以及基于虚假信息与虚假差评数据训练的模型(DIS+NEG)。这四个混合模型检测欺骗与虚假信息的预测准确率在75%至85%之间。我们使用SHAP对模型结果进行评估,以确定各特征的影响程度。