Defeasibility in causal reasoning implies that the causal relationship between cause and effect can be strengthened or weakened. Namely, the causal strength between cause and effect should increase or decrease with the incorporation of strengthening arguments (supporters) or weakening arguments (defeaters), respectively. However, existing works ignore defeasibility in causal reasoning and fail to evaluate existing causal strength metrics in defeasible settings. In this work, we present {\delta}-CAUSAL, the first benchmark dataset for studying defeasibility in causal reasoning. {\delta}-CAUSAL includes around 11K events spanning ten domains, featuring defeasible causality pairs, i.e., cause-effect pairs accompanied by supporters and defeaters. We further show current causal strength metrics fail to reflect the change of causal strength with the incorporation of supporters or defeaters in {\delta}-CAUSAL. To this end, we propose CESAR (Causal Embedding aSsociation with Attention Rating), a metric that measures causal strength based on token-level causal relationships. CESAR achieves a significant 69.7% relative improvement over existing metrics, increasing from 47.2% to 80.1% in capturing the causal strength change brought by supporters and defeaters. We further demonstrate even Large Language Models (LLMs) like GPT-3.5 still lag 4.5 and 10.7 points behind humans in generating supporters and defeaters, emphasizing the challenge posed by {\delta}-CAUSAL.
翻译:因果推理中的可废止性意味着因果效应与原因之间的因果关系可以增强或减弱。即,原因与效应之间的因果强度应随着强化论证(支持者)或弱化论证(反对者)的加入而分别增加或减少。然而,现有研究忽略了因果推理中的可废止性,未能评估现有因果强度度量在可废止场景下的表现。本文提出了δ-CAUSAL,这是首个用于研究因果推理中可废止性的基准数据集。δ-CAUSAL包含跨越十个领域的约1.1万个事件,具有可废止因果对,即附带支持者和反对者的因果-效应对。我们进一步表明,现有因果强度度量无法反映δ-CAUSAL中因果强度随支持者或反对者加入的变化。为此,我们提出CESAR(基于注意力评分的因果嵌入关联度量),这是一种基于词级因果关系测量因果强度的度量。CESAR在捕捉支持者和反对者带来的因果强度变化方面,相比现有度量实现了69.7%的相对提升,从47.2%提升至80.1%。我们进一步证明,即使是GPT-3.5这样的大型语言模型(LLM)在生成支持者和反对者方面仍落后人类4.5和10.7个百分点,强调了δ-CAUSAL带来的挑战。