Interest in understanding and factorizing learned embedding spaces through conceptual explanations is steadily growing. When no human concept labels are available, concept discovery methods search trained embedding spaces for interpretable concepts like object shape or color that can provide post-hoc explanations for decisions. Unlike previous work, we argue that concept discovery should be identifiable, meaning that a number of known concepts can be provably recovered to guarantee reliability of the explanations. As a starting point, we explicitly make the connection between concept discovery and classical methods like Principal Component Analysis and Independent Component Analysis by showing that they can recover independent concepts under non-Gaussian distributions. For dependent concepts, we propose two novel approaches that exploit functional compositionality properties of image-generating processes. Our provably identifiable concept discovery methods substantially outperform competitors on a battery of experiments including hundreds of trained models and dependent concepts, where they exhibit up to 29 % better alignment with the ground truth. Our results highlight the strict conditions under which reliable concept discovery without human labels can be guaranteed and provide a formal foundation for the domain. Our code is available online.
翻译:对通过学习得到的嵌入空间进行理解与因子化并借助概念解释的兴趣正日益增长。当缺乏人类概念标签时,概念发现方法在训练后的嵌入空间中搜索可解释的概念(如物体形状或颜色),从而为决策提供后验解释。与以往工作不同,我们认为概念发现应具备可识别性,即能够确保可靠地恢复一定数量的已知概念,以保证解释的可靠性。作为起点,我们明确建立了概念发现与主成分分析、独立成分分析等经典方法之间的联系,证明它们能够在非高斯分布下恢复独立概念。针对依赖概念,我们提出了两种利用图像生成过程功能组合性特征的新方法。我们的可证明可识别的概念发现方法在包含数百个训练模型与依赖概念的一系列实验中显著优于现有方法,其与真实概念的匹配度最高提升达29%。本研究揭示了在没有人类标签的情况下保证可靠概念发现的严格条件,并为本领域提供了形式化基础。相关代码已公开。