The explosive growth of online content demands robust Natural Language Processing (NLP) techniques that can capture nuanced meanings and cultural context across diverse languages. Semantic Textual Relatedness (STR) goes beyond superficial word overlap, considering linguistic elements and non-linguistic factors like topic, sentiment, and perspective. Despite its pivotal role, prior NLP research has predominantly focused on English, limiting its applicability across languages. Addressing this gap, our paper dives into capturing deeper connections between sentences beyond simple word overlap. Going beyond English-centric NLP research, we explore STR in Marathi, Hindi, Spanish, and English, unlocking the potential for information retrieval, machine translation, and more. Leveraging the SemEval-2024 shared task, we explore various language models across three learning paradigms: supervised, unsupervised, and cross-lingual. Our comprehensive methodology gains promising results, demonstrating the effectiveness of our approach. This work aims to not only showcase our achievements but also inspire further research in multilingual STR, particularly for low-resourced languages.
翻译:在线内容的爆炸式增长对自然语言处理技术提出了更高要求,需能在不同语言中捕捉细微语义及文化语境。语义文本相关性超越了表层词汇重叠,考虑了语言要素及话题、情感、视角等非语言因素。尽管其作用关键,但先前的自然语言处理研究主要聚焦英语,限制了跨语言适用性。为填补这一空白,本文致力于捕捉句子间超越简单词汇重叠的深层关联。突破以英语为中心的自然语言处理研究范式,我们探索了马拉地语、印地语、西班牙语和英语的语义文本相关性,为信息检索、机器翻译等领域释放潜力。借助SemEval-2024共享任务,我们分别在监督学习、无监督学习及跨语言学习三种范式下考察了多种语言模型。该综合性方法取得了显著成效,验证了其有效性。本研究不仅旨在展示成果,更希望激发对多语言语义文本相关性(尤其是低资源语言领域)的进一步探索。