Large Language Models (LLMs) have been progressively exhibiting there capabilities in various areas of research. The performance of the LLMs in acute maternal healthcare area, predominantly in low resource languages like Telugu, Hindi, Tamil, Urdu etc are still unstudied. This study presents how ChatGPT-4o, GeminiAI, and Perplexity AI respond to pregnancy related questions asked in different languages. A bilingual dataset is used to obtain results by applying the semantic similarity metrics (BERT Score) and expert assessments from expertise gynecologists. Multiple parameters like accuracy, fluency, relevance, coherence and completeness are taken into consideration by the gynecologists to rate the responses generated by the LLMs. Gemini excels in other LLMs in terms of producing accurate and coherent pregnancy relevant responses in Telugu, while Perplexity demonstrated well when the prompts were in Telugu. ChatGPT's performance can be improved. The results states that both selecting an LLM and prompting language plays a crucial role in retrieving the information. Altogether, we emphasize for the improvement of LLMs assistance in regional languages for healthcare purposes.
翻译:大型语言模型(LLMs)在多个研究领域逐渐展现出其能力。然而,这些模型在急性母婴保健领域(尤其是泰卢固语、印地语、泰米尔语、乌尔都语等低资源语言)的性能仍未得到充分研究。本研究展示了ChatGPT-4o、GeminiAI和Perplexity AI如何回应以不同语言提出的妊娠相关问题。研究采用双语数据集,通过语义相似度指标(BERT Score)以及经验丰富的妇科专家的专业评估来获取结果。妇科专家综合考虑准确性、流畅性、相关性、连贯性和完整性等多个参数,对LLMs生成的回复进行评分。在生成准确且连贯的泰卢固语妊娠相关回复方面,Gemini优于其他LLMs,而Perplexity在提示词为泰卢固语时表现出色。ChatGPT的性能仍有提升空间。结果表明,LLM的选择和提示语言在信息检索中均起着关键作用。总体而言,我们强调需要改进LLMs在区域语言中对医疗保健领域的辅助功能。