Despite their success, Large-Language Models (LLMs) still face criticism as their lack of interpretability limits their controllability and reliability. Traditional post-hoc interpretation methods, based on attention and gradient-based analysis, offer limited insight into the model's decision-making processes. In the image field, Concept-based models have emerged as explainable-by-design architectures, employing human-interpretable features as intermediate representations. However, these methods have not been yet adapted to textual data, mainly because they require expensive concept annotations, which are impractical for real-world text data. This paper addresses this challenge by proposing a self-supervised Interpretable Concept Embedding Models (ICEMs). We leverage the generalization abilities of LLMs to predict the concepts labels in a self-supervised way, while we deliver the final predictions with an interpretable function. The results of our experiments show that ICEMs can be trained in a self-supervised way achieving similar performance to fully supervised concept-based models and end-to-end black-box ones. Additionally, we show that our models are (i) interpretable, offering meaningful logical explanations for their predictions; (ii) interactable, allowing humans to modify intermediate predictions through concept interventions; and (iii) controllable, guiding the LLMs' decoding process to follow a required decision-making path.
翻译:尽管大型语言模型(LLM)取得了成功,但其可解释性的缺乏限制了其可控性与可靠性,因而仍面临批评。传统的基于注意力与梯度的后验解释方法,对模型决策过程的洞察力有限。在图像领域,基于概念的模型作为一种可解释性设计架构出现,其采用人类可解释的特征作为中间表示。然而,这些方法尚未被成功应用于文本数据,主要原因是它们需要昂贵的概念标注,而这对于现实世界的文本数据而言并不实用。本文通过提出一种自监督的可解释概念嵌入模型(ICEM)来应对这一挑战。我们利用LLM的泛化能力以自监督方式预测概念标签,同时通过一个可解释函数给出最终预测。我们的实验结果表明,ICEM可以通过自监督方式进行训练,并达到与全监督基于概念的模型以及端到端黑盒模型相当的性能。此外,我们证明我们的模型具有以下特性:(i)可解释性,能够为其预测提供有意义的逻辑解释;(ii)可交互性,允许人类通过概念干预来修改中间预测;(iii)可控性,能够引导LLM的解码过程遵循所需的决策路径。