The quality of explanations for the predictions of complex machine learning predictors is often measured using insertion and deletion metrics, which assess the faithfulness of the explanations, i.e., how correctly the explanations reflect the predictor's behavior. To improve the faithfulness, we propose insertion/deletion metric-aware explanation-based optimization (ID-ExpO), which optimizes differentiable predictors to improve both insertion and deletion scores of the explanations while keeping their predictive accuracy. Since the original insertion and deletion metrics are indifferentiable with respect to the explanations and directly unavailable for gradient-based optimization, we extend the metrics to be differentiable and use them to formalize insertion and deletion metric-based regularizers. The experimental results on image and tabular datasets show that the deep neural networks-based predictors fine-tuned using ID-ExpO enable popular post-hoc explainers to produce more faithful and easy-to-interpret explanations while keeping high predictive accuracy.
翻译:对于复杂机器学习模型预测结果的解释质量,通常通过插入和删除度量进行评估,这些度量衡量解释的忠实度,即解释正确反映预测器行为的程度。为提高忠实度,我们提出插入/删除度量感知的解释性优化(ID-ExpO),该方法在保持预测准确性的同时,通过优化可微预测器来改善解释的插入与删除分数。由于原始插入和删除度量关于解释不可微,无法直接用于基于梯度的优化,我们将这些度量扩展为可微形式,并据此形式化插入与删除度量正则化器。在图像和表格数据集上的实验结果表明,采用ID-ExpO微调的深度神经网络预测器,能够使主流事后解释器在保持高预测准确性的同时,生成更忠实且易于解释的解释结果。