The ability to explain why a machine learning model arrives at a particular prediction is crucial when used as decision support by human operators of critical systems. The provided explanations must be provably correct, and preferably without redundant information, called minimal explanations. In this paper, we aim at finding explanations for predictions made by tree ensembles that are not only minimal, but also minimum with respect to a cost function. To this end, we first present a highly efficient oracle that can determine the correctness of explanations, surpassing the runtime performance of current state-of-the-art alternatives by several orders of magnitude when computing minimal explanations. Secondly, we adapt an algorithm called MARCO from related works (calling it m-MARCO) for the purpose of computing a single minimum explanation per prediction, and demonstrate an overall speedup factor of two compared to the MARCO algorithm which enumerates all minimal explanations. Finally, we study the obtained explanations from a range of use cases, leading to further insights of their characteristics. In particular, we observe that in several cases, there are more than 100,000 minimal explanations to choose from for a single prediction. In these cases, we see that only a small portion of the minimal explanations are also minimum, and that the minimum explanations are significantly less verbose, hence motivating the aim of this work.
翻译:在为关键系统提供决策支持时,机器学习模型为何得出特定预测的解释能力至关重要。所提供的解释必须可证明正确,并且最好不包含冗余信息,即所谓的最小化解释。本文旨在为树集成模型做出的预测寻找解释,这些解释不仅是最小化的,而且相对于代价函数而言也是最小代价的。为此,我们首先提出一种高效的可满足性检查器,能够判定解释的正确性,在计算最小化解释时,其运行性能比当前最先进的替代方案高出数个数量级。其次,我们改编了相关工作中的MARCO算法(称之为m-MARCO),用于为每个预测计算单一的最小代价解释,并展示了相较于枚举所有最小化解释的MARCO算法,整体加速比达到两倍。最后,我们通过一系列用例研究了所获得的解释,进一步洞悉了它们的特性。特别地,我们观察到,在多种情况下,单个预测有超过10万个最小化解释可供选择。在这些情况下,只有极小一部分最小化解释同时也是最小代价解释,且最小代价解释的冗长度显著降低,这进一步印证了本文研究工作的动机。