Siamese networks have gained popularity as a method for modeling text semantic similarity. Traditional methods rely on pooling operation to compress the semantic representations from Transformer blocks in encoding, resulting in two-dimensional semantic vectors and the loss of hierarchical semantic information from Transformer blocks. Moreover, this limited structure of semantic vectors is akin to a flattened landscape, which restricts the methods that can be applied in downstream modeling, as they can only navigate this flat terrain. To address this issue, we propose a novel 3D Siamese network for text semantic similarity modeling, which maps semantic information to a higher-dimensional space. The three-dimensional semantic tensors not only retains more precise spatial and feature domain information but also provides the necessary structural condition for comprehensive downstream modeling strategies to capture them. Leveraging this structural advantage, we introduce several modules to reinforce this 3D framework, focusing on three aspects: feature extraction, attention, and feature fusion. Our extensive experiments on four text semantic similarity benchmarks demonstrate the effectiveness and efficiency of our 3D Siamese Network.
翻译:孪生网络作为一种文本语义相似度建模方法已受到广泛关注。传统方法依赖池化操作压缩编码器中Transformer块生成的语义表征,导致产生二维语义向量并丢失Transformer块中的层次化语义信息。此外,这种受限的语义向量结构如同扁平化景观,限制了可应用于下游建模的方法——它们仅能在此平坦地形中导航。针对此问题,我们提出了一种新颖的三维孪生网络用于文本语义相似度建模,该方法将语义信息映射至更高维空间。三维语义张量不仅保留了更精确的空间与特征域信息,还为全面的下游建模策略捕捉这些信息提供了必要的结构条件。借助这一结构优势,我们从特征提取、注意力机制与特征融合三方面引入若干模块以强化该三维框架。在四个文本语义相似度基准上的大量实验表明,我们的三维孪生网络具有显著的有效性与高效性。