Question Answering (QA) datasets have been instrumental in developing and evaluating Large Language Model (LLM) capabilities. However, such datasets are scarce for languages other than English due to the cost and difficulties of collection and manual annotation. This means that producing novel models and measuring the performance of multilingual LLMs in low-resource languages is challenging. To mitigate this, we propose $\textbf{S}$yn$\textbf{DAR}$in, a method for generating and validating QA datasets for low-resource languages. We utilize parallel content mining to obtain $\textit{human-curated}$ paragraphs between English and the target language. We use the English data as context to $\textit{generate}$ synthetic multiple-choice (MC) question-answer pairs, which are automatically translated and further validated for quality. Combining these with their designated non-English $\textit{human-curated}$ paragraphs form the final QA dataset. The method allows to maintain the content quality, reduces the likelihood of factual errors, and circumvents the need for costly annotation. To test the method, we created a QA dataset with $1.2$K samples for the Armenian language. The human evaluation shows that $98\%$ of the generated English data maintains quality and diversity in the question types and topics, while the translation validation pipeline can filter out $\sim70\%$ of data with poor quality. We use the dataset to benchmark state-of-the-art LLMs, showing their inability to achieve human accuracy with some model performances closer to random chance. This shows that the generated dataset is non-trivial and can be used to evaluate reasoning capabilities in low-resource language.
翻译:问答(QA)数据集对于开发和评估大语言模型(LLM)的能力至关重要。然而,由于收集和人工标注的成本与难度,除英语外的其他语言普遍缺乏此类数据集。这意味着为低资源语言开发新模型或评估多语言LLM的性能面临挑战。为缓解这一问题,我们提出$\textbf{S}$yn$\textbf{DAR}$in方法,用于为低资源语言生成和验证QA数据集。我们利用平行内容挖掘技术获取英语与目标语言之间$\textit{人工校订}$的段落,并以英文数据作为上下文$\textit{生成}$合成的多项选择(MC)问答对,再通过自动翻译并进行质量验证。将这些合成数据与对应的非英语$\textit{人工校订}$段落结合,即构成最终的QA数据集。该方法能够保持内容质量,降低事实错误的可能性,并规避昂贵的人工标注成本。为验证该方法,我们为亚美尼亚语构建了包含$1.2$K样本的QA数据集。人工评估表明,生成的英文数据中$98\%$在问题类型和主题上保持了质量与多样性,而翻译验证流程可过滤约$70\%$的低质量数据。我们使用该数据集对前沿LLM进行基准测试,结果显示这些模型均无法达到人类准确率,部分模型的性能甚至接近随机猜测。这表明所生成的数据集具有实质性难度,可用于评估低资源语言中的推理能力。