Despite significant progress, evaluation of explainable artificial intelligence remains elusive and challenging. In this paper we propose a fine-grained validation framework that is not overly reliant on any one facet of these sociotechnical systems, and that recognises their inherent modular structure: technical building blocks, user-facing explanatory artefacts and social communication protocols. While we concur that user studies are invaluable in assessing the quality and effectiveness of explanation presentation and delivery strategies from the explainees' perspective in a particular deployment context, the underlying explanation generation mechanisms require a separate, predominantly algorithmic validation strategy that accounts for the technical and human-centred desiderata of their (numerical) outputs. Such a comprehensive sociotechnical utility-based evaluation framework could allow to systematically reason about the properties and downstream influence of different building blocks from which explainable artificial intelligence systems are composed -- accounting for a diverse range of their engineering and social aspects -- in view of the anticipated use case.
翻译:尽管取得了显著进展,解释型人工智能的评估仍然难以捉摸且充满挑战。本文提出一种细粒度验证框架,该框架不过度依赖这些社会技术系统的单一维度,同时认识到其固有的模块化结构:技术构建模块、面向用户的解释性产物及社会沟通协议。虽然我们认同用户研究在特定部署场景中,从解释接收者视角评估解释呈现与传递策略的质量与有效性方面具有不可替代的价值,但底层的解释生成机制需要采用独立的、以算法为主的验证策略,以兼顾其(数值)输出的技术需求与以人为本的设计准则。这种全面的社会技术效用导向评估框架,能够系统性地推理构成解释型人工智能系统的不同模块的属性与下游影响——涵盖其工程与社会层面的多元维度——并针对预期应用场景进行考量。