Interpretable representations are the backbone of many explainers that target black-box predictive systems based on artificial intelligence and machine learning algorithms. They translate the low-level data representation necessary for good predictive performance into high-level human-intelligible concepts used to convey the explanatory insights. Notably, the explanation type and its cognitive complexity are directly controlled by the interpretable representation, tweaking which allows to target a particular audience and use case. However, many explainers built upon interpretable representations overlook their merit and fall back on default solutions that often carry implicit assumptions, thereby degrading the explanatory power and reliability of such techniques. To address this problem, we study properties of interpretable representations that encode presence and absence of human-comprehensible concepts. We demonstrate how they are operationalised for tabular, image and text data; discuss their assumptions, strengths and weaknesses; identify their core building blocks; and scrutinise their configuration and parameterisation. In particular, this in-depth analysis allows us to pinpoint their explanatory properties, desiderata and scope for (malicious) manipulation in the context of tabular data where a linear model is used to quantify the influence of interpretable concepts on a black-box prediction. Our findings lead to a range of recommendations for designing trustworthy interpretable representations; specifically, the benefits of class-aware (supervised) discretisation of tabular data, e.g., with decision trees, and sensitivity of image interpretable representations to segmentation granularity and occlusion colour.
翻译:可解释表征是许多针对基于人工智能和机器学习算法的黑箱预测系统的解释器的核心支柱。它们将实现良好预测性能所需的低层数据表征转化为用于传达解释性洞察的高层人类可理解概念。值得注意的是,解释类型及其认知复杂度直接受可解释表征控制,调整该表征可针对特定受众和使用场景。然而,许多基于可解释表征构建的解释器忽视了其价值,退而采用往往隐含假设的默认解决方案,从而削弱了此类技术的解释能力和可靠性。为解决该问题,本研究探讨了编码人类可理解概念存在与否的可解释表征属性。我们演示了它们如何用于表格、图像和文本数据;讨论其假设、优缺点;识别其核心构建模块;并审视其配置与参数化。特别地,这一深度分析使我们能够精准定位其在表格数据场景中的解释属性、期望特征及(恶意)操纵空间——此场景下使用线性模型量化可解释概念对黑箱预测的影响。研究结果形成了一套设计可信赖可解释表征的建议;具体包括:表格数据类别感知(监督式)离散化(如使用决策树)的优势,以及图像可解释表征对分割粒度和遮挡颜色的敏感性。