Explainable artificial intelligence (XAI) is motivated by the problem of making AI predictions understandable, transparent, and responsible, as AI becomes increasingly impactful in society and high-stakes domains. XAI algorithms are designed to explain AI decisions in human-understandable ways. The evaluation and optimization criteria of XAI are gatekeepers for XAI algorithms to achieve their expected goals and should withstand rigorous inspection. To improve the scientific rigor of XAI, we conduct the first critical examination of a common XAI criterion: plausibility. It measures how convincing the AI explanation is to humans, and is usually quantified by metrics on feature localization or correlation of feature attribution. Our examination shows, although plausible explanations can improve users' understanding and local trust in an AI decision, doing so is at the cost of abandoning other possible approaches of enhancing understandability, increasing misleading explanations that manipulate users, being unable to achieve complementary human-AI task performance, and deteriorating users' global trust in the overall AI system. Because the flaws outweigh the benefits, we do not recommend using plausibility as a criterion to evaluate or optimize XAI algorithms. We also identify new directions to improve XAI on understandability and utility to users including complementary human-AI task performance.
翻译:随着人工智能在社会和高风险领域的影响力日益增强,可解释人工智能(XAI)旨在使人工智能的预测变得可理解、透明且负责任。XAI算法被设计用于以人类可理解的方式解释人工智能的决策。XAI的评估与优化标准是确保算法实现预期目标的关键门槛,必须经得起严格检验。为提升XAI的科学严谨性,本文首次对一项常用的XAI评估标准——可理解性——展开批判性审视。该标准通过特征定位或特征归因相关性的度量指标,衡量人工智能解释对人类的说服力。研究表明,尽管可理解的解释能提升用户对特定AI决策的理解与局部信任,但其代价包括:放弃其他提升可理解性的可能途径、增加误导性解释操纵用户的风险、无法实现人机协同的任务性能,以及削弱用户对整体AI系统的全局信任。由于弊大于利,我们不建议将可理解性作为评估或优化XAI算法的标准。同时,本文提出了改进XAI在可理解性与用户效用(包括人机协同任务性能)方面的新方向。