Attention mechanisms have played a crucial role in the development of complex architectures such as Transformers in natural language processing. However, Transformers remain hard to interpret and are considered as black-boxes. This paper aims to assess how attention coefficients from Transformers can help in providing interpretability. A new attention-based interpretability method called CLaSsification-Attention (CLS-A) is proposed. CLS-A computes an interpretability score for each word based on the attention coefficient distribution related to the part specific to the classification task within the Transformer architecture. A human-grounded experiment is conducted to evaluate and compare CLS-A to other interpretability methods. The experimental protocol relies on the capacity of an interpretability method to provide explanation in line with human reasoning. Experiment design includes measuring reaction times and correct response rates by human subjects. CLS-A performs comparably to usual interpretability methods regarding average participant reaction time and accuracy. The lower computational cost of CLS-A compared to other interpretability methods and its availability by design within the classifier make it particularly interesting. Data analysis also highlights the link between the probability score of a classifier prediction and adequate explanations. Finally, our work confirms the relevancy of the use of CLS-A and shows to which extent self-attention contains rich information to explain Transformer classifiers.
翻译:注意力机制在自然语言处理中复杂架构(如Transformer)的发展中发挥了关键作用。然而,Transformer仍然难以解释,被视为"黑箱"。本文旨在评估Transformer中的注意力系数如何帮助提升可解释性。我们提出了一种名为分类-注意力(CLS-A)的新型基于注意力的可解释性方法。CLS-A通过分析Transformer架构中与分类任务特定部分相关的注意力系数分布,为每个单词计算可解释性分数。通过一项基于人类实验的研究,我们对CLS-A与其他可解释性方法进行了评估与比较。该实验协议依赖于可解释性方法是否能够提供符合人类推理的解释。实验设计包括测量人类受试者的反应时和正确应答率。在参与者平均反应时间和准确率方面,CLS-A的表现与常规可解释性方法相当。相较于其他可解释性方法,CLS-A具有较低的计算成本,并且其作为分类器设计中的固有属性而可直接使用,这使其特别引人注目。数据分析还揭示了分类器预测的概率分数与充分解释之间的关联。最后,我们的研究证实了CLS-A的相关性,并展示了自注意力机制在多大程度上包含解释Transformer分类器的丰富信息。