A common approach to quantifying model interpretability is to calculate faithfulness metrics based on iteratively masking input tokens and measuring how much the predicted label changes as a result. However, we show that such metrics are generally not suitable for comparing the interpretability of different neural text classifiers as the response to masked inputs is highly model-specific. We demonstrate that iterative masking can produce large variation in faithfulness scores between comparable models, and show that masked samples are frequently outside the distribution seen during training. We further investigate the impact of adversarial attacks and adversarial training on faithfulness scores, and demonstrate the relevance of faithfulness measures for analyzing feature salience in text adversarial attacks. Our findings provide new insights into the limitations of current faithfulness metrics and key considerations to utilize them appropriately.
翻译:衡量模型可解释性的一种常见方法是基于迭代掩码输入标记并测量预测标签相应变化程度的忠实性指标。然而,我们表明这类指标通常不适用于比较不同神经文本分类器的可解释性,因为对掩码输入的响应高度依赖于具体模型。我们证明迭代掩码可能会导致可比模型间的忠实性得分存在较大差异,并表明掩码样本常常处于训练时未见的数据分布之外。我们进一步研究了对抗攻击和对抗训练对忠实性得分的影响,并证明了忠实性度量在分析文本对抗攻击中特征显著性方面的相关性。我们的发现为当前忠实性指标的局限性提供了新见解,并提出了恰当使用这些指标的关键考量因素。