Food classification is an important task in health care. In this work, we propose a multimodal classification framework that uses the modified version of EfficientNet with the Mish activation function for image classification, and the traditional BERT transformer-based network is used for text classification. The proposed network and the other state-of-the-art methods are evaluated on a large open-source dataset, UPMC Food-101. The experimental results show that the proposed network outperforms the other methods, a significant difference of 11.57% and 6.34% in accuracy is observed for image and text classification, respectively, when compared with the second-best performing method. We also compared the performance in terms of accuracy, precision, and recall for text classification using both machine learning and deep learning-based models. The comparative analysis from the prediction results of both images and text demonstrated the efficiency and robustness of the proposed approach.
翻译:食物分类在医疗健康领域是一项重要任务。本文提出了一种多模态分类框架,该框架采用改进版EfficientNet(结合Mish激活函数)进行图像分类,并利用传统基于BERT Transformer的网络进行文本分类。所提出的网络及其他先进方法在大型开源数据集UPMC Food-101上进行了评估。实验结果表明,所提出网络优于其他方法:与次优方法相比,图像分类和文本分类的准确率分别显著提高了11.57%和6.34%。我们还使用基于机器学习和深度学习的模型,从准确率、精确率和召回率三个维度对文本分类性能进行了比较。基于图像和文本预测结果的对比分析证实了所提出方法的有效性和鲁棒性。