Current evaluation benchmarks for question answering (QA) in Indic languages often rely on machine translation of existing English datasets. This approach suffers from bias and inaccuracies inherent in machine translation, leading to datasets that may not reflect the true capabilities of EQA models for Indic languages. This paper proposes a new benchmark specifically designed for evaluating Hindi EQA models and discusses the methodology to do the same for any task. This method leverages large language models (LLMs) to generate a high-quality dataset in an extractive setting, ensuring its relevance for the target language. We believe this new resource will foster advancements in Hindi NLP research by providing a more accurate and reliable evaluation tool.
翻译:当前用于印度语言问答评估的基准数据集往往依赖于对现有英语数据集的机器翻译。这种方法存在机器翻译固有的偏差和不准确性,导致生成的数据集可能无法真实反映印度语言EQA模型的实际能力。本文提出了一项专为评估印地语EQA模型设计的新基准,并讨论了将该方法推广至任意任务的技术路线。该方法利用大型语言模型在抽取式场景中生成高质量数据集,确保其与目标语言的关联性。我们相信这一新资源将通过提供更精确可靠的评估工具,推动印地语自然语言处理研究的发展。