In the wake of responsible AI, interpretability methods, which attempt to provide an explanation for the predictions of neural models have seen rapid progress. In this work, we are concerned with explanations that are applicable to natural language processing (NLP) models and tasks, and we focus specifically on the analysis of counterfactual, contrastive explanations. We note that while there have been several explainers proposed to produce counterfactual explanations, their behaviour can vary significantly and the lack of a universal ground truth for the counterfactual edits imposes an insuperable barrier on their evaluation. We propose a new back translation-inspired evaluation methodology that utilises earlier outputs of the explainer as ground truth proxies to investigate the consistency of explainers. We show that by iteratively feeding the counterfactual to the explainer we can obtain valuable insights into the behaviour of both the predictor and the explainer models, and infer patterns that would be otherwise obscured. Using this methodology, we conduct a thorough analysis and propose a novel metric to evaluate the consistency of counterfactual generation approaches with different characteristics across available performance indicators.
翻译:在负责任的AI浪潮推动下,旨在为神经网络模型预测提供解释的可解释性方法取得了快速发展。本研究聚焦于适用于自然语言处理模型与任务的解释方法,特别关注反事实对比性解释的分析。我们注意到,尽管已有多种解释器被提出用于生成反事实解释,但其行为差异显著,且缺乏统一的反事实编辑真实基准,这对其评估构成了难以逾越的障碍。我们提出了一种受反向翻译启发的新型评估方法,该方法利用解释器早期输出作为真实基准代理,以探究解释器的一致性。实验表明,通过将反事实结果迭代输入解释器,我们能够获得关于预测模型与解释模型行为的重要洞见,并推断出其他方法难以发现的模式。基于这一方法论,我们进行了全面分析,并提出了一个新颖的度量标准,用于评估不同特征的反事实生成方法在现有性能指标上的一致性。