We evaluate the ability of large language models (LLMs) to infer causal relations from natural language. Compared to traditional natural language processing and deep learning techniques, LLMs show competitive performance in a benchmark of pairwise relations without needing (explicit) training samples. This motivates us to extend our approach to extrapolating causal graphs through iterated pairwise queries. We perform a preliminary analysis on a benchmark of biomedical abstracts with ground-truth causal graphs validated by experts. The results are promising and support the adoption of LLMs for such a crucial step in causal inference, especially in medical domains, where the amount of scientific text to analyse might be huge, and the causal statements are often implicit.
翻译:我们评估了大语言模型从自然语言中推断因果关系的能力。相较于传统自然语言处理与深度学习技术,大语言模型在无需(显式)训练样本的条件下,于成对关系基准测试中展现出具有竞争力的性能。这促使我们将方法扩展至通过迭代成对查询外推因果图。我们利用经专家验证的基准数据集(包含真实因果关系的生物医学摘要)进行了初步分析。实验结果表明,大语言模型在因果推断这一关键环节具有应用潜力,尤其适用于医学领域——该领域需要分析的海量科学文本中因果陈述往往隐含其中。