Setting proper evaluation objectives for explainable artificial intelligence (XAI) is vital for making XAI algorithms follow human communication norms, support human reasoning processes, and fulfill human needs for AI explanations. In this article, we examine explanation plausibility, which is the most pervasive human-grounded concept in XAI evaluation. Plausibility measures how reasonable the machine explanation is compared to the human explanation. Plausibility has been conventionally formulated as an important evaluation objective for AI explainability tasks. We argue against this idea, and show how optimizing and evaluating XAI for plausibility is sometimes harmful, and always ineffective to achieve model understandability, transparency, and trustworthiness. Specifically, evaluating XAI algorithms for plausibility regularizes the machine explanation to express exactly the same content as human explanation, which deviates from the fundamental motivation for humans to explain: expressing similar or alternative reasoning trajectories while conforming to understandable forms or language. Optimizing XAI for plausibility regardless of the model decision correctness also jeopardizes model trustworthiness, as doing so breaks an important assumption in human-human explanation namely that plausible explanations typically imply correct decisions, and violating this assumption eventually leads to either undertrust or overtrust of AI models. Instead of being the end goal in XAI evaluation, plausibility can serve as an intermediate computational proxy for the human process of interpreting explanations to optimize the utility of XAI. We further highlight the importance of explainability-specific evaluation objectives by differentiating the AI explanation task from the object localization task.
翻译:为可解释人工智能(XAI)设定适当的评估目标,对于使XAI算法遵循人类交流规范、支持人类推理过程并满足人类对AI解释的需求至关重要。本文审视了解释的合理性——这是XAI评估中最普遍的人类基础概念。合理性衡量机器解释与人类解释相比的合理程度。传统上,合理性被定义为AI可解释性任务的重要评估目标。我们对此提出反对,并论证了优化和评估XAI的合理性有时有害,且始终无法有效实现模型的可理解性、透明度和可信度。具体而言,以合理性为目标评估XAI算法会使机器解释表达与人类解释完全相同的内容,这偏离了人类解释的根本动机:以可理解的形式或语言表达相似或替代的推理轨迹。不考虑模型决策正确性而优化XAI的合理性也会损害模型可信度,因为这打破了人类解释中的一个重要假设——合理的解释通常意味着正确的决策,而违反此假设最终会导致对AI模型的过度不信任或过度信任。合理性不应是XAI评估的最终目标,而可作为人类解释过程的中介计算代理,以优化XAI的实用性。我们进一步通过区分AI解释任务与对象定位任务,强调了可解释性专属评估目标的重要性。