Interest in understanding and factorizing learned embedding spaces through conceptual explanations is steadily growing. When no human concept labels are available, concept discovery methods search trained embedding spaces for interpretable concepts like object shape or color that can be used to provide post-hoc explanations for decisions. Unlike previous work, we argue that concept discovery should be identifiable, meaning that a number of known concepts can be provably recovered to guarantee reliability of the explanations. As a starting point, we explicitly make the connection between concept discovery and classical methods like Principal Component Analysis and Independent Component Analysis by showing that they can recover independent concepts with non-Gaussian distributions. For dependent concepts, we propose two novel approaches that exploit functional compositionality properties of image-generating processes. Our provably identifiable concept discovery methods substantially outperform competitors on a battery of experiments including hundreds of trained models and dependent concepts, where they exhibit up to 29 % better alignment with the ground truth. Our results provide a rigorous foundation for reliable concept discovery without human labels.
翻译:对通过学习得到的嵌入空间进行理解与分解的兴趣正日益增长。当缺乏人工概念标签时,概念发现方法在已训练的嵌入空间中搜索可解释的概念(如物体形状或颜色),这些概念可用于为决策提供后验解释。与以往工作不同,我们认为概念发现应具备可识别性,即能够可证明地恢复若干已知概念,从而保障解释的可靠性。作为起点,我们明确建立了概念发现与主成分分析和独立成分分析等经典方法之间的联系,证明这些方法可以恢复具有非高斯分布的独立概念。针对依赖概念,我们提出了两种新方法,利用图像生成过程的函数组合性质。在涵盖数百个训练模型及依赖概念的系列实验中,我们的可证明可识别概念发现方法显著优于竞争方法,与真实标注的匹配度提升高达29%。我们的结果为无需人工标签的可靠概念发现提供了严谨的理论基础。