Prior work has demonstrated large language models' (LLMs) potential to discern statistical tendencies within their pre-training corpora. Despite that, many examinations of LLMs' knowledge capacity focus on knowledge explicitly appearing in the training data or implicitly inferable from similar contexts. How well an LLM captures the corpus-level statistical trends of concepts for reasoning, especially long-tail ones, is still underexplored. In this study, we introduce a novel few-shot question-answering task (CPopQA) that examines LLMs' statistical ranking abilities for long-tail cultural concepts (e.g., holidays), with a specific focus on these concepts' popularity in the United States and the United Kingdom, respectively. We curate a dataset containing 459 holidays across 58 countries, generating a total of 6,000 QA testing pairs. Experiments on four strong LLMs show that large models are capable of ranking long-tail cultural concepts regarding their statistical tendency. Notably, GPT-3.5 displayed superior performance and exhibited its potential to identify geo-cultural proximity across continents.
翻译:先前研究已证明大语言模型(LLMs)具备感知其预训练语料库中统计倾向的潜力。尽管如此,许多关于LLMs知识能力的考察仍聚焦于训练数据中显式出现或可从相似语境隐式推断的知识。对于LLMs捕捉语料库层面概念推理(尤其是长尾概念)统计趋势的能力,目前研究尚不充分。本研究引入一项新颖的少样本问答任务(CPopQA),旨在检验LLMs对长尾文化概念(如节假日)的统计排名能力,重点关注这些概念在美国和英国的流行度差异。我们整理了一个涵盖58个国家共459个节假日的数据集,生成了总计6000组问答测试对。在四个强LLMs上的实验表明,大型模型能够对长尾文化概念的统计倾向进行排名。值得注意的是,GPT-3.5展现出卓越性能,并表现出识别跨大洲地理文化接近性的潜力。