Despite the rapid development of large language models (LLMs) for the Korean language, there remains an obvious lack of benchmark datasets that test the requisite Korean cultural and linguistic knowledge. Because many existing Korean benchmark datasets are derived from the English counterparts through translation, they often overlook the different cultural contexts. For the few benchmark datasets that are sourced from Korean data capturing cultural knowledge, only narrow tasks such as bias and hate speech detection are offered. To address this gap, we introduce a benchmark of Cultural and Linguistic Intelligence in Korean (CLIcK), a dataset comprising 1,995 QA pairs. CLIcK sources its data from official Korean exams and textbooks, partitioning the questions into eleven categories under the two main categories of language and culture. For each instance in CLIcK, we provide fine-grained annotation of which cultural and linguistic knowledge is required to answer the question correctly. Using CLIcK, we test 13 language models to assess their performance. Our evaluation uncovers insights into their performances across the categories, as well as the diverse factors affecting their comprehension. CLIcK offers the first large-scale comprehensive Korean-centric analysis of LLMs' proficiency in Korean culture and language.
翻译:尽管针对韩语的大型语言模型(LLMs)发展迅速,但测试韩语必要文化与语言知识的基准数据集明显匮乏。由于许多现有韩语基准数据集是通过翻译英语对应数据集得到的,它们常常忽视不同的文化背景。即便有少数从捕捉文化知识的韩语数据中构建的基准数据集,也仅提供偏见与仇恨言论检测等狭隘任务。为填补这一空白,我们引入韩语文化与语言智能基准数据集CLIcK,该数据集包含1,995个问答对。CLIcK从官方韩语考试与教科书中收集数据,将问题划分为语言与文化两大类下的共十一个子类别。对于CLIcK中的每个实例,我们提供细粒度标注,明确回答该问题所需的文化与语言知识。我们利用CLIcK测试了13种语言模型以评估其性能。评估结果揭示了模型在各类别上的表现差异,以及影响其理解能力的多种因素。CLIcK首次提供了面向韩语的大规模综合性LLMs韩语文化与语言能力分析。