Evaluating the accuracy of outputs generated by Large Language Models (LLMs) is especially important in the climate science and policy domain. We introduce the Expert Confidence in Climate Statements (ClimateX) dataset, a novel, curated, expert-labeled dataset consisting of 8094 climate statements collected from the latest Intergovernmental Panel on Climate Change (IPCC) reports, labeled with their associated confidence levels. Using this dataset, we show that recent LLMs can classify human expert confidence in climate-related statements, especially in a few-shot learning setting, but with limited (up to 47%) accuracy. Overall, models exhibit consistent and significant over-confidence on low and medium confidence statements. We highlight implications of our results for climate communication, LLMs evaluation strategies, and the use of LLMs in information retrieval systems.
翻译:评估大型语言模型(LLMs)生成输出的准确性在气候科学与政策领域尤为重要。我们引入了气候声明专家置信度数据集(ClimateX),这是一个经过精心策划、专家标注的新数据集,包含来自最新政府间气候变化专门委员会(IPCC)报告的8094条气候声明,并标注了相应的置信度水平。利用该数据集,我们展示了近期LLMs能够对气候相关声明中人类专家的置信度进行分类,尤其是在少样本学习设置下,但准确率有限(最高达47%)。总体而言,模型在低置信度和中等置信度的声明上表现出持续且显著的过度自信。我们强调了这些结果对气候传播、LLMs评估策略以及LLMs在信息检索系统中应用的意义。