Words of estimative probability (WEP) are expressions of a statement's plausibility (probably, maybe, likely, doubt, likely, unlikely, impossible...). Multiple surveys demonstrate the agreement of human evaluators when assigning numerical probability levels to WEP. For example, highly likely corresponds to a median chance of 0.90+-0.08 in Fagen-Ulmschneider (2015)'s survey. In this work, we measure the ability of neural language processing models to capture the consensual probability level associated to each WEP. Firstly, we use the UNLI dataset (Chen et al., 2020) which associates premises and hypotheses with their perceived joint probability p, to construct prompts, e.g. "[PREMISE]. [WEP], [HYPOTHESIS]." and assess whether language models can predict whether the WEP consensual probability level is close to p. Secondly, we construct a dataset of WEP-based probabilistic reasoning, to test whether language models can reason with WEP compositions. When prompted "[EVENTA] is likely. [EVENTB] is impossible.", a causal language model should not express that [EVENTA&B] is likely. We show that both tasks are unsolved by off-the-shelf English language models, but that fine-tuning leads to transferable improvement.
翻译:推测性概率词汇(WEP)是表述陈述可信度的词语(如很可能、也许、大概、怀疑、不太可能、不可能等)。多项调查表明,人类评估者在将数值概率水平分配给WEP时具有一致性。例如,在Fagen-Ulmschneider(2015)的调查中,“高度可能”对应的中位数概率为0.90±0.08。本研究衡量神经语言处理模型捕捉每个WEP所对应的共识概率水平的能力。首先,我们利用UNLI数据集(Chen等,2020),该数据集将前提与假设与其感知联合概率p相关联,构建提示模板如“[前提]。[WEP],[假设]。”并评估语言模型能否预测WEP的共识概率水平是否接近p。其次,我们构建基于WEP的概率推理数据集,以测试语言模型能否处理WEP组合推理。例如,当提示“[事件A]很可能。[事件B]不可能”时,因果语言模型不应表达“[事件A与B]很可能”这一结论。我们证明,现有英文语言模型在两个任务中均表现不佳,但微调可带来可迁移的改进。