Over the past few years, different types of data-driven Artificial Intelligence (AI) techniques have been widely adopted in various domains of science for generating predictive models. However, because of their black-box nature, it is crucial to establish trust in these models before accepting them as accurate. One way of achieving this goal is through the implementation of a post-hoc interpretation scheme that can put forward the reasons behind a black-box model's prediction. In this work, we propose a classical thermodynamics inspired approach for this purpose: Thermodynamically Explainable Representations of AI and other black-box Paradigms (TERP). TERP works by constructing a linear, local surrogate model that approximates the behaviour of the black-box model within a small neighborhood around the instance being explained. By employing a simple forward feature selection algorithm, TERP assigns an interpretability score to all the possible surrogate models. Compared to existing methods, TERP improves interpretability by selecting an optimal interpretation from these models by drawing simple parallels with classical thermodynamics. To validate TERP as a generally applicable method, we successfully demonstrate how it can be used to obtain interpretations of a wide range of black-box model architectures including deep learning Autoencoders, Recurrent neural networks and Convolutional neural networks applied to different domains including molecular simulations, image, and text classification respectively.
翻译:近年来,不同类型的数据驱动型人工智能技术已被广泛应用于各科学领域以生成预测模型。然而,由于其黑箱特性,在将这些模型认定为准确之前,建立对其的信任至关重要。实现这一目标的一种方法是实施事后解释方案,该方案能够阐明黑箱模型预测背后的原因。本文为此提出一种受经典热力学启发的方法:人工智能及其他黑箱范例的热力学可解释表示。该方法通过构建一个线性局部代理模型来运作,该模型在待解释实例的小邻域内近似黑箱模型的行为。通过采用简单的前向特征选择算法,对所有可能的代理模型赋予可解释性评分。与现有方法相比,通过借鉴经典热力学的简单类比,从这些模型中选择最优解释,从而提升了可解释性。为验证TERP作为通用方法的有效性,我们成功展示了如何将其应用于获得多种黑箱模型架构的解释,包括应用于分子模拟、图像分类和文本分类等不同领域的深度学习自编码器、递归神经网络和卷积神经网络。