Counterfactual explanations (CFEs) are a popular approach in explainable artificial intelligence (xAI), highlighting changes to input data necessary for altering a model's output. A CFE can either describe a scenario that is better than the factual state (upward CFE), or a scenario that is worse than the factual state (downward CFE). However, potential benefits and drawbacks of the directionality of CFEs for user behavior in xAI remain unclear. The current user study (N=161) compares the impact of CFE directionality on behavior and experience of participants tasked to extract new knowledge from an automated system based on model predictions and CFEs. Results suggest that upward CFEs provide a significant performance advantage over other forms of counterfactual feedback. Moreover, the study highlights potential benefits of mixed CFEs improving user performance compared to downward CFEs or no explanations. In line with the performance results, users' explicit knowledge of the system is statistically higher after receiving upward CFEs compared to downward comparisons. These findings imply that the alignment between explanation and task at hand, the so-called regulatory fit, may play a crucial role in determining the effectiveness of model explanations, informing future research directions in xAI. To ensure reproducible research, the entire code, underlying models and user data of this study is openly available: https://github.com/ukuhl/DirectionalAlienZoo
翻译:反事实解释(CFEs)是可解释人工智能(xAI)中的一种流行方法,它突出显示改变模型输出所需的输入数据变化。一个CFE可以描述比事实状态更好的场景(向上CFE),或者比事实状态更差的场景(向下CFE)。然而,CFE方向性对xAI用户行为的潜在益处与缺点仍不明确。本用户研究(N=161)比较了CFE方向性对行为与体验的影响,参与者需基于模型预测和CFE从自动化系统中提取新知识。结果表明,向上CFE相比其他形式的反事实反馈提供了显著的性能优势。此外,研究强调了混合CFE在提升用户性能方面相较于向下CFE或无解释的潜在益处。与性能结果一致,接受向上CFE后用户对系统的显性知识显著高于接受向下比较的情况。这些发现表明,解释与手头任务的契合度(即所谓的调节匹配)可能在决定模型解释有效性中起关键作用,为xAI的未来研究方向提供参考。为确保研究可复现,本研究的全部代码、底层模型和用户数据已公开提供:https://github.com/ukuhl/DirectionalAlienZoo