Many ethical frameworks require artificial intelligence (AI) systems to be explainable. Explainable AI (XAI) models are frequently tested for their adequacy in user studies. Since different people may have different explanatory needs, it is important that participant samples in user studies are large enough to represent the target population to enable generalizations. However, it is unclear to what extent XAI researchers reflect on and justify their sample sizes or avoid broad generalizations across people. We analyzed XAI user studies (N = 220) published between 2012 and 2022. Most studies did not offer rationales for their sample sizes. Moreover, most papers generalized their conclusions beyond their target population, and there was no evidence that broader conclusions in quantitative studies were correlated with larger samples. These methodological problems can impede evaluations of whether XAI systems implement the explainability called for in ethical frameworks. We outline principles for more inclusive XAI user studies.
翻译:许多伦理框架要求人工智能系统具备可解释性。可解释人工智能模型通常通过用户研究来测试其充分性。由于不同个体存在不同的解释需求,用户研究的参与者样本需具有足够规模以代表目标人群,从而支持结论泛化。然而,目前尚不清楚可解释人工智能领域研究者对样本量的反思与论证程度,以及他们在结论中是否避免过度泛化。我们分析了2012年至2022年间发表的220项可解释人工智能用户研究,发现大多数研究未对样本量选择提供充分理由,且多数论文将其结论泛化至目标人群之外,而量化研究的结论泛化范围与样本规模之间未呈现显著相关性。这些方法论问题可能阻碍评估可解释人工智能系统是否真正实现伦理框架所要求的可解释性。为此,我们提出更具包容性的可解释人工智能用户研究设计原则。