Model interpretability plays a central role in human-AI decision-making systems. Ideally, explanations should be expressed using human-interpretable semantic concepts. Moreover, the causal relations between these concepts should be captured by the explainer to allow for reasoning about the explanations. Lastly, explanation methods should be efficient and not compromise the performance of the predictive task. Despite the rapid advances in AI explainability in recent years, as far as we know to date, no method fulfills these three properties. Indeed, mainstream methods for local concept explainability do not produce causal explanations and incur a trade-off between explainability and prediction performance. We present DiConStruct, an explanation method that is both concept-based and causal, with the goal of creating more interpretable local explanations in the form of structural causal models and concept attributions. Our explainer works as a distillation model to any black-box machine learning model by approximating its predictions while producing the respective explanations. Because of this, DiConStruct generates explanations efficiently while not impacting the black-box prediction task. We validate our method on an image dataset and a tabular dataset, showing that DiConStruct approximates the black-box models with higher fidelity than other concept explainability baselines, while providing explanations that include the causal relations between the concepts.
翻译:模型可解释性在人机决策系统中扮演核心角色。理想情况下,解释应使用人类可理解的语义概念来表达。此外,解释器应捕捉这些概念之间的因果关系,以便进行关于解释的推理。最后,解释方法应高效且不影响预测任务的性能。尽管近年来AI可解释性取得了快速进展,但据我们所知,目前尚无方法同时满足这三个特性。事实上,主流局部概念可解释性方法无法生成因果解释,且会在可解释性与预测性能之间产生权衡。我们提出DiConStruct,一种基于概念且具备因果性的解释方法,旨在以结构因果模型和概念归因的形式创建更可解释的局部解释。我们的解释器作为蒸馏模型,通过近似任意黑箱机器学习模型的预测并生成相应解释来工作。因此,DiConStruct高效生成解释且不影响黑箱预测任务。我们在图像数据集和表格数据集上验证了该方法,结果表明DiConStruct在近似黑箱模型时比其它概念可解释性基线具有更高的保真度,同时提供包含概念间因果关系的解释。