This work introduces interpretable regional descriptors, or IRDs, for local, model-agnostic interpretations. IRDs are hyperboxes that describe how an observation's feature values can be changed without affecting its prediction. They justify a prediction by providing a set of "even if" arguments (semi-factual explanations), and they indicate which features affect a prediction and whether pointwise biases or implausibilities exist. A concrete use case shows that this is valuable for both machine learning modelers and persons subject to a decision. We formalize the search for IRDs as an optimization problem and introduce a unifying framework for computing IRDs that covers desiderata, initialization techniques, and a post-processing method. We show how existing hyperbox methods can be adapted to fit into this unified framework. A benchmark study compares the methods based on several quality measures and identifies two strategies to improve IRDs.
翻译:本文引入了可解释的区域描述符(Interpretable Regional Descriptors,简称IRDs),用于局部、模型无关的解释。IRDs是描述观测特征值在不影响其预测结果的前提下可以如何变化的超盒。它们通过提供一组“即便”论证(半事实解释)来证明预测的合理性,并指明哪些特征影响预测,以及是否存在逐点偏差或不可信之处。一个具体用例表明,这对机器学习建模者和受决策影响的个人都具有价值。我们将IRDs的搜索形式化为一个优化问题,并引入一个统一的IRDs计算框架,该框架涵盖了期望特性、初始化技术以及后处理方法。我们展示了如何使现有的超盒方法适应这一统一框架。一项基准研究基于若干质量指标对多种方法进行了比较,并识别出两种改进IRDs的策略。