Setting proper evaluation objectives for explainable artificial intelligence (XAI) is vital for making XAI algorithms follow human communication norms, support human reasoning processes, and fulfill human needs for AI explanations. In this position paper, we examine the most pervasive human-grounded concept in XAI evaluation, explanation plausibility. Plausibility measures how reasonable the machine explanation is compared to the human explanation. Plausibility has been conventionally formulated as an important evaluation objective for AI explainability tasks. We argue against this idea, and show how optimizing and evaluating XAI for plausibility is sometimes harmful, and always ineffective in achieving model understandability, transparency, and trustworthiness. Specifically, evaluating XAI algorithms for plausibility regularizes the machine explanation to express exactly the same content as human explanation, which deviates from the fundamental motivation for humans to explain: expressing similar or alternative reasoning trajectories while conforming to understandable forms or language. Optimizing XAI for plausibility regardless of the model decision correctness also jeopardizes model trustworthiness, because doing so breaks an important assumption in human-human explanation that plausible explanations typically imply correct decisions, and vice versa; and violating this assumption eventually leads to either undertrust or overtrust of AI models. Instead of being the end goal in XAI evaluation, plausibility can serve as an intermediate computational proxy for the human process of interpreting explanations to optimize the utility of XAI. We further highlight the importance of explainability-specific evaluation objectives by differentiating the AI explanation task from the object localization task.
翻译:为可解释人工智能(XAI)设定合适的评估目标,对于使XAI算法遵循人类沟通规范、支持人类推理过程并满足人类对AI解释的需求至关重要。在这篇立场论文中,我们审视了XAI评估中最普遍的人类基础概念——解释合理性(plausibility)。合理性衡量机器解释与人类解释相比的合理程度。传统上,合理性被设定为AI可解释性任务的重要评估目标。我们对此观点提出质疑,并论证了以合理性为目标优化和评估XAI有时是有害的,且始终无法有效实现模型的可理解性、透明度和可信赖性。具体而言,以合理性为标准评估XAI算法会迫使机器解释表达与人类解释完全一致的内容,这偏离了人类进行解释的根本动机——在遵循可理解形式或语言的同时,表达相似或替代性的推理轨迹。无视模型决策正确性而优化XAI的合理性也会损害模型可信赖性,因为这打破了人-人解释中的一个重要假设:合理的解释通常暗示正确的决策(反之亦然);违反这一假设最终会导致对AI模型的信任不足或过度信任。我们进一步指出,合理性不应成为XAI评估的最终目标,而可作为人类解释理解过程的中介计算代理,用于优化XAI的效用。通过区分AI解释任务与目标定位任务,我们强调了以可解释性为导向的评估目标的重要性。