Ensuring the trustworthiness and interpretability of machine learning models is critical to their deployment in real-world applications. Feature attribution methods have gained significant attention, which provide local explanations of model predictions by attributing importance to individual input features. This study examines the generalization of feature attributions across various deep learning architectures, such as convolutional neural networks (CNNs) and vision transformers. We aim to assess the feasibility of utilizing a feature attribution method as a future detector and examine how these features can be harmonized across multiple models employing distinct architectures but trained on the same data distribution. By exploring this harmonization, we aim to develop a more coherent and optimistic understanding of feature attributions, enhancing the consistency of local explanations across diverse deep-learning models. Our findings highlight the potential for harmonized feature attribution methods to improve interpretability and foster trust in machine learning applications, regardless of the underlying architecture.
翻译:确保机器学习模型的可信度和可解释性对其在现实应用中的部署至关重要。特征归因方法通过将重要性赋予单个输入特征,为模型预测提供局部解释,因而受到广泛关注。本研究考察了特征归因在多种深度学习架构(如卷积神经网络和视觉变换器)中的泛化能力。我们旨在评估将特征归因方法用作未来检测器的可行性,并探讨如何在使用不同架构但基于相同数据分布训练的多个模型中协调这些特征。通过探索这种协调,我们旨在形成对特征归因更为一致且乐观的理解,从而增强不同深度学习模型之间局部解释的一致性。我们的研究结果突显了协调化特征归因方法在改进可解释性及提升机器学习应用可信度方面的潜力,不受底层架构影响。