At the staggering pace with which the capabilities of large language models (LLMs) are increasing, creating future-proof evaluation sets to assess their understanding becomes more and more challenging. In this paper, we propose a novel paradigm for evaluating LLMs which leverages the idea that correct world understanding should be consistent across different (Fregean) senses of the same meaning. Accordingly, we measure understanding not in terms of correctness but by evaluating consistency across multiple senses that are generated by the model itself. We showcase our approach by instantiating a test where the different senses are different languages, hence using multilingual self-consistency as a litmus test for the model's understanding and simultaneously addressing the important topic of multilingualism. Taking one of the latest versions of ChatGPT as our object of study, we evaluate multilingual consistency for two different tasks across three different languages. We show that its multilingual consistency is still lacking, and that its task and world understanding are thus not language-independent. As our approach does not require any static evaluation corpora in languages other than English, it can easily and cheaply be extended to different languages and tasks and could become an integral part of future benchmarking efforts.
翻译:随着大型语言模型(LLMs)能力的飞速提升,构建经得起未来考验的评估集以衡量其理解能力变得愈发具有挑战性。本文提出一种新颖的LLM评估范式,其核心理念是:对世界的正确理解应在同一含义的不同(弗雷格)意义上保持一致。据此,我们不通过正确性来度量理解,而是评估模型自身生成的多种意义间的一致性。我们通过实例化一项测试来展示该方法:将不同的意义设置为不同语言,从而以多语言自一致性作为模型理解的试金石,同时探讨多语言这一重要议题。以最新版本的ChatGPT为研究对象,我们评估了其跨三种语言、两项任务的多语言一致性。结果表明,其多语言一致性仍有欠缺,因此其任务与世界的理解并非语言无关。由于该方法无需任何非英语的静态评估语料库,因此能够轻松且低成本地扩展到不同语言和任务,并有望成为未来基准测试工作的核心组成部分。