This paper presents a novel concept learning framework for enhancing model interpretability and performance in visual classification tasks. Our approach appends an unsupervised explanation generator to the primary classifier network and makes use of adversarial training. During training, the explanation module is optimized to extract visual concepts from the classifier's latent representations, while the GAN-based module aims to discriminate images generated from concepts, from true images. This joint training scheme enables the model to implicitly align its internally learned concepts with human-interpretable visual properties. Comprehensive experiments demonstrate the robustness of our approach, while producing coherent concept activations. We analyse the learned concepts, showing their semantic concordance with object parts and visual attributes. We also study how perturbations in the adversarial training protocol impact both classification and concept acquisition. In summary, this work presents a significant step towards building inherently interpretable deep vision models with task-aligned concept representations - a key enabler for developing trustworthy AI for real-world perception tasks.
翻译:本文提出了一种新颖的概念学习框架,旨在增强视觉分类任务中模型的可解释性与性能。我们的方法在主分类器网络后附加一个无监督解释生成器,并采用对抗训练。训练过程中,解释模块被优化以从分类器的潜在表示中提取视觉概念,而基于GAN的模块则致力于区分由概念生成的图像与真实图像。这种联合训练方案使模型能够隐式地将其内部学习的概念与人类可解释的视觉属性对齐。综合实验证明了我们方法的鲁棒性,同时产生了一致的概念激活模式。我们分析了所学概念,展示了它们与物体部件及视觉属性在语义上的一致性。我们还研究了对抗训练协议中的扰动如何影响分类与概念获取。总之,这项工作朝着构建具有任务对齐概念表示的内在可解释深度视觉模型迈出了重要一步——这是为现实世界感知任务开发可信人工智能的关键推动因素。