An evaluation of LLMs and Google Translate for translation of selected Indian languages via sentiment and semantic analyses

Large Language models (LLMs) have been prominent for language translation, including low-resource languages. There has been limited study about the assessment of the quality of translations generated by LLMs, including Gemini, GPT and Google Translate. In this study, we address this limitation by using semantic and sentiment analysis of selected LLMs for Indian languages, including Sanskrit, Telugu and Hindi. We select prominent texts that have been well translated by experts and use LLMs to generate their translations to English, and then we provide a comparison with selected expert (human) translations. Our findings suggest that while LLMs have made significant progress in translation accuracy, challenges remain in preserving sentiment and semantic integrity, especially in figurative and philosophical contexts. The sentiment analysis revealed that GPT-4o and GPT-3.5 are better at preserving the sentiments for the Bhagavad Gita (Sanskrit-English) translations when compared to Google Translate. We observed a similar trend for the case of Tamas (Hindi-English) and Maha P (Telugu-English) translations. GPT-4o performs similarly to GPT-3.5 in the translation in terms of sentiments for the three languages. We found that LLMs are generally better at translation for capturing sentiments when compared to Google Translate.

翻译：大型语言模型（LLM）在语言翻译领域表现突出，包括低资源语言翻译。然而，针对LLM（如Gemini、GPT系列）及谷歌翻译生成译文质量的评估研究仍较为有限。本研究通过语义分析与情感分析，评估了LLM在梵语、泰卢固语和印地语等选定印度语言翻译中的表现。我们选取专家优质译本的经典文本，使用LLM将其译为英文，并与专家（人工）译本进行对比。研究发现：尽管LLM在翻译准确性方面取得显著进展，但在保持情感与语义完整性方面仍存在挑战，尤其在比喻性与哲学性语境中。情感分析表明，在《薄伽梵歌》（梵语-英语）翻译中，GPT-4o与GPT-3.5在情感保留方面优于谷歌翻译。在《Tamas》（印地语-英语）与《Maha P》（泰卢固语-英语）翻译中也观察到相似趋势。就三种语言翻译的情感保持而言，GPT-4o与GPT-3.5表现相当。总体而言，LLM在翻译中捕捉情感的能力普遍优于谷歌翻译。

相关内容

Google

关注 77

一家美国的跨国科技企业，致力于互联网搜索、云计算、广告技术等领域，由当时在斯坦福大学攻读理学博士的拉里·佩奇和谢尔盖·布林共同创建。创始之初，Google 官方的公司使命为「整合全球范围的信息，使人人皆可访问并从中受益」。 Google 开发并提供了大量基于互联网的产品与服务，其主要利润来自于 AdWords 等广告服务。

2004 年 8 月 19 日，公司以「GOOG」为代码正式登陆纳斯达克交易所。

Linux导论，Introduction to Linux，96页ppt

专知会员服务

82+阅读 · 2020年7月26日

【跨语言BERT模型大集合】Transfer learning is increasingly going multilingual with language-specific BERT models

专知会员服务

54+阅读 · 2020年1月30日

FlowQA: Grasping Flow in History for Conversational Machine Comprehension

专知会员服务

34+阅读 · 2019年10月18日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

50+阅读 · 2019年10月17日