Determining the degree of confidence of deep learning model in its prediction is an open problem in the field of natural language processing. Most of the classical methods for uncertainty estimation are quite weak for text classification models. We set the task of obtaining an uncertainty estimate for neural networks based on the Transformer architecture. A key feature of such mo-dels is the attention mechanism, which supports the information flow between the hidden representations of tokens in the neural network. We explore the formed relationships between internal representations using Topological Data Analysis methods and utilize them to predict model's confidence. In this paper, we propose a method for uncertainty estimation based on the topological properties of the attention mechanism and compare it with classical methods. As a result, the proposed algorithm surpasses the existing methods in quality and opens up a new area of application of the attention mechanism, but requires the selection of topological features.
翻译:确定深度学习模型在其预测中的置信度是自然语言处理领域的一个开放性问题。大多数经典的不确定性估计方法对于文本分类模型效果较弱。我们设定了为基于Transformer架构的神经网络获取不确定性估计的任务。这类模型的一个关键特征是注意力机制,它支持神经网络中词元隐藏表示之间的信息流动。我们利用拓扑数据分析方法探索内部表示之间形成的关系,并利用这些关系来预测模型的置信度。在本文中,我们提出了一种基于注意力机制拓扑性质的不确定性估计方法,并将其与经典方法进行了比较。结果表明,所提出的算法在质量上超越了现有方法,开辟了注意力机制的新应用领域,但需要选择拓扑特征。