In this paper, we propose a novel approach for generating document embeddings using a combination of Sentence-BERT (SBERT) and RoBERTa, two state-of-the-art natural language processing models. Our approach treats sentences as tokens and generates embeddings for them, allowing the model to capture both intra-sentence and inter-sentence relations within a document. We evaluate our model on a book recommendation task and demonstrate its effectiveness in generating more semantically rich and accurate document embeddings. To assess the performance of our approach, we conducted experiments on a book recommendation task using the Goodreads dataset. We compared the document embeddings generated using our MULTI-BERT model to those generated using SBERT alone. We used precision as our evaluation metric to compare the quality of the generated embeddings. Our results showed that our model consistently outperformed SBERT in terms of the quality of the generated embeddings. Furthermore, we found that our model was able to capture more nuanced semantic relations within documents, leading to more accurate recommendations. Overall, our results demonstrate the effectiveness of our approach and suggest that it is a promising direction for improving the performance of recommendation systems
翻译:本文提出了一种利用Sentence-BERT(SBERT)与RoBERTa两种先进自然语言处理模型组合生成文档嵌入的新方法。该方法将句子视为标记并为其生成嵌入,使模型能够捕捉文档内句子内部及句子之间的语义关系。我们在书籍推荐任务上评估了模型性能,证明其能生成语义更丰富、更准确的文档嵌入。为评估方法效果,我们使用Goodreads数据集开展了书籍推荐实验,将基于MULTI-BERT模型生成的文档嵌入与仅使用SBERT生成的嵌入进行对比,并以精确率作为评估指标衡量嵌入质量。结果表明,本模型在生成嵌入质量上始终优于SBERT。此外,我们发现该模型能捕捉文档中更细微的语义关系,从而产生更精准的推荐结果。总体而言,实验结果验证了方法的有效性,表明其为提升推荐系统性能提供了有前景的研究方向。