Model explanations are very valuable for interpreting and debugging prediction models. We study a specific kind of global explanations called Concept Explanations, where the goal is to interpret a model using human-understandable concepts. Recent advances in multi-modal learning rekindled interest in concept explanations and led to several label-efficient proposals for estimation. However, existing estimation methods are unstable to the choice of concepts or dataset that is used for computing explanations. We observe that instability in explanations is due to high variance in point estimation of importance scores. We propose an uncertainty aware Bayesian estimation method, which readily improved reliability of the concept explanations. We demonstrate with theoretical analysis and empirical evaluation that explanations computed by our method are more reliable while also being label-efficient and faithful.
翻译:模型解释对于解读和调试预测模型非常有价值。我们研究一种称为概念解释的全局解释类型,其目标是使用人类可理解的概念来解释模型。多模态学习的最新进展重新点燃了人们对概念解释的兴趣,并推动了多种节省标注的评估方法。然而,现有的评估方法在用于计算解释的概念或数据集的选择上存在不稳定性。我们观察到,解释的不稳定性源于重要度分数点估计的高方差。我们提出了一种考虑不确定性的贝叶斯估计方法,该方法显著提高了概念解释的可靠性。通过理论分析和实证评估,我们证明使用我们的方法计算出的解释更具可靠性,同时兼具标签高效性和忠实性。