Deep learning methods are highly accurate, yet their opaque decision process prevents them from earning full human trust. Concept-based models aim to address this issue by learning tasks based on a set of human-understandable concepts. However, state-of-the-art concept-based models rely on high-dimensional concept embedding representations which lack a clear semantic meaning, thus questioning the interpretability of their decision process. To overcome this limitation, we propose the Deep Concept Reasoner (DCR), the first interpretable concept-based model that builds upon concept embeddings. In DCR, neural networks do not make task predictions directly, but they build syntactic rule structures using concept embeddings. DCR then executes these rules on meaningful concept truth degrees to provide a final interpretable and semantically-consistent prediction in a differentiable manner. Our experiments show that DCR: (i) improves up to +25% w.r.t. state-of-the-art interpretable concept-based models on challenging benchmarks (ii) discovers meaningful logic rules matching known ground truths even in the absence of concept supervision during training, and (iii), facilitates the generation of counterfactual examples providing the learnt rules as guidance.
翻译:深度学习方法具有高精度,然而其不透明的决策过程阻碍了其获得人类的完全信任。基于概念的模型旨在通过基于一组人类可理解的概念来学习任务以解决这一问题。然而,最先进的基于概念的模型依赖于缺乏清晰语义含义的高维概念嵌入表示,从而使其决策过程的可解释性受到质疑。为克服这一局限,我们提出了深度概念推理器(DCR),这是首个基于概念嵌入的可解释概念模型。在DCR中,神经网络不直接进行任务预测,而是利用概念嵌入构建句法规则结构。随后,DCR在具有意义的概念真值度上执行这些规则,以可微分的方式提供最终可解释且语义一致的预测。我们的实验表明,DCR:(i)在具有挑战性的基准测试上,相较于最先进的可解释概念模型性能提升高达+25%;(ii)即使在训练期间缺乏概念监督的情况下,也能发现与已知真实规则相匹配的有意义的逻辑规则;(iii)便于生成反事实示例,并将学得的规则作为指导。