Semantic similarity analysis and modeling is a fundamentally acclaimed task in many pioneering applications of natural language processing today. Owing to the sensation of sequential pattern recognition, many neural networks like RNNs and LSTMs have achieved satisfactory results in semantic similarity modeling. However, these solutions are considered inefficient due to their inability to process information in a non-sequential manner, thus leading to the improper extraction of context. Transformers function as the state-of-the-art architecture due to their advantages like non-sequential data processing and self-attention. In this paper, we perform semantic similarity analysis and modeling on the U.S Patent Phrase to Phrase Matching Dataset using both traditional and transformer-based techniques. We experiment upon four different variants of the Decoding Enhanced BERT - DeBERTa and enhance its performance by performing K-Fold Cross-Validation. The experimental results demonstrate our methodology's enhanced performance compared to traditional techniques, with an average Pearson correlation score of 0.79.
翻译:语义相似性分析与建模是当今自然语言处理众多前沿应用中的一项关键基础任务。受序列模式识别的启发,RNN和LSTM等神经网络已在语义相似性建模方面取得了令人满意的成果。然而,这些方法因其无法以非序列方式处理信息而导致上下文提取不当,因此被认为效率低下。Transformer凭借其非序列数据处理和自注意力机制等优势,成为当前最先进的架构。本文使用传统技术和基于Transformer的技术,对美国专利短语间匹配数据集进行语义相似性分析与建模。我们对解码增强型BERT——DeBERTa的四种不同变体进行了实验,并通过执行K折交叉验证来提升其性能。实验结果表明,与传统技术相比,我们的方法性能更优,平均皮尔逊相关系数达到0.79。