In the interest of interpreting neural NLI models and their reasoning strategies, we carry out a systematic probing study which investigates whether these models capture the crucial semantic features central to natural logic: monotonicity and concept inclusion. Correctly identifying valid inferences in downward-monotone contexts is a known stumbling block for NLI performance, subsuming linguistic phenomena such as negation scope and generalized quantifiers. To understand this difficulty, we emphasize monotonicity as a property of a context and examine the extent to which models capture monotonicity information in the contextual embeddings which are intermediate to their decision making process. Drawing on the recent advancement of the probing paradigm, we compare the presence of monotonicity features across various models. We find that monotonicity information is notably weak in the representations of popular NLI models which achieve high scores on benchmarks, and observe that previous improvements to these models based on fine-tuning strategies have introduced stronger monotonicity features together with their improved performance on challenge sets.
翻译:为探究神经NLI模型的推理策略及其可解释性,我们开展了一项系统性探测研究,重点考察这些模型是否捕捉到自然逻辑中的核心语义特征:单调性与概念包含关系。在向下单调语境中正确识别有效推理是NLI性能提升的关键障碍,该问题涵盖否定辖域和广义量词等语言现象。为深入理解这一难点,我们将单调性视为语境属性,通过分析模型决策过程中间层生成的上下文嵌入,评估模型对单调性信息的捕捉程度。基于探测范式的最新进展,我们比较了不同模型中单调性特征的存在情况。研究发现,在基准测试中表现优异的主流NLI模型的表征中,单调性信息显著薄弱;同时观察到,基于微调策略的模型改进在提升挑战集性能的同时,也引入了更强的单调性特征。