Deep machine learning models including Convolutional Neural Networks (CNN) have been successful in the detection of Mild Cognitive Impairment (MCI) using medical images, questionnaires, and videos. This paper proposes a novel Multi-branch Classifier-Video Vision Transformer (MC-ViViT) model to distinguish MCI from those with normal cognition by analyzing facial features. The data comes from the I-CONECT, a behavioral intervention trial aimed at improving cognitive function by providing frequent video chats. MC-ViViT extracts spatiotemporal features of videos in one branch and augments representations by the MC module. The I-CONECT dataset is challenging as the dataset is imbalanced containing Hard-Easy and Positive-Negative samples, which impedes the performance of MC-ViViT. We propose a loss function for Hard-Easy and Positive-Negative Samples (HP Loss) by combining Focal loss and AD-CORRE loss to address the imbalanced problem. Our experimental results on the I-CONECT dataset show the great potential of MC-ViViT in predicting MCI with a high accuracy of 90.63\% accuracy on some of the interview videos.
翻译:包括卷积神经网络(CNN)在内的深度机器学习模型已成功应用于通过医学图像、问卷和视频检测轻度认知障碍(MCI)。本文提出了一种新型多分支分类器-视频视觉Transformer(MC-ViViT)模型,通过分析面部特征区分MCI患者与认知正常老年人。数据来源于I-CONECT行为干预试验,该试验通过提供频繁视频通话改善认知功能。MC-ViViT在一个分支中提取视频的时空特征,并通过MC模块增强表征。I-CONECT数据集具有挑战性,因存在类别不平衡现象,包含难易样本(Hard-Easy)与正负样本(Positive-Negative),这制约了MC-ViViT的性能。我们针对难易样本与正负样本提出结合Focal损失和AD-CORRE损失的HP损失函数,以解决不平衡问题。在I-CONECT数据集上的实验结果表明,MC-ViViT在部分访谈视频上预测MCI的准确率高达90.63%,展现出巨大潜力。