Explainable Artificial Intelligence (XAI) aims to make learning machines less opaque, and offers researchers and practitioners various tools to reveal the decision-making strategies of neural networks. In this work, we investigate how XAI methods can be used for exploring and visualizing the diversity of feature representations learned by Bayesian Neural Networks (BNNs). Our goal is to provide a global understanding of BNNs by making their decision-making strategies a) visible and tangible through feature visualizations and b) quantitatively measurable with a distance measure learned by contrastive learning. Our work provides new insights into the \emph{posterior} distribution in terms of human-understandable feature information with regard to the underlying decision making strategies. The main findings of our work are the following: 1) global XAI methods can be applied to explain the diversity of decision-making strategies of BNN instances, 2) Monte Carlo dropout with commonly used Dropout rates exhibit increased diversity in feature representations compared to the multimodal posterior approximation of MultiSWAG, 3) the diversity of learned feature representations highly correlates with the uncertainty estimate for the output and 4) the inter-mode diversity of the multimodal posterior decreases as the network width increases, while the intra mode diversity increases. These findings are consistent with the recent Deep Neural Networks theory, providing additional intuitions about what the theory implies in terms of humanly understandable concepts.
翻译:可解释人工智能(XAI)旨在减少学习机器的不透明性,并为研究人员和从业者提供各种工具来揭示神经网络的决策策略。本文探究了如何利用XAI方法探索并可视化贝叶斯神经网络(BNNs)学习到的特征表示多样性。我们的目标是通过以下方式提供对BNNs的全局理解:a) 通过特征可视化使决策策略可见且可感知,b) 使用对比学习得到的距离度量进行定量测量。本文从人类可理解的特征信息角度,针对底层决策策略,为理解后验分布提供了新见解。主要发现如下:1) 全局XAI方法可用于解释BNN实例决策策略的多样性;2) 与MultiSWAG的多模态后验近似相比,采用常用丢弃率的蒙特卡洛丢弃法在特征表示中表现出更高的多样性;3) 学习到的特征表示多样性与输出的不确定性估计高度相关;4) 随着网络宽度增加,多模态后验的模态间多样性降低,而模态内多样性增加。这些发现与近期深度神经网络理论一致,为该理论在人类可理解概念层面的含义提供了额外直观理解。