In current NLP research, large-scale language models and their abilities are widely being discussed. Some recent works have also found notable failures of these models. Often these failure examples involve complex reasoning abilities. This work focuses on a simple commonsense ability, reasoning about when an action (or its effect) is feasible. To this end, we introduce FeasibilityQA, a question-answering dataset involving binary classification (BCQ) and multi-choice multi-correct questions (MCQ) that test understanding of feasibility. We show that even state-of-the-art models such as GPT-3, GPT-2, and T5 struggle to answer the feasibility questions correctly. Specifically, on MCQ and BCQ questions, GPT-3 achieves an accuracy of just (19%, 62%) and (25%, 64%) in zero-shot and few-shot settings, respectively. We also evaluate models by providing relevant knowledge statements required to answer the question. We find that the additional knowledge leads to a 7% gain in performance, but the overall performance still remains low. These results make one wonder how much commonsense knowledge about action feasibility is encoded in state-of-the-art models and how well they can reason about it.
翻译:在当前NLP研究中,大规模语言模型及其能力被广泛探讨。近期研究也发现了这些模型的显著缺陷,这些失败案例通常涉及复杂推理能力。本研究聚焦于一项简单的常识能力:对行动(或其效果)可行性的推理。为此,我们提出FeasibilityQA——一个包含二元分类题(BCQ)和多选题(MCQ)的问答数据集,用于测试对可行性的理解。实验表明,即便是GPT-3、GPT-2和T5等最新模型,在正确回答可行性问题上仍存在困难。具体而言,在零样本和少样本设置下,GPT-3在MCQ与BCQ问题上的准确率仅分别达到(19%、62%)和(25%、64%)。我们还通过提供回答问题所需的相关知识陈述来评估模型,发现额外知识使性能提升了7%,但整体表现仍处于较低水平。这些结果令人质疑:关于行动可行性的常识知识在最新模型中被编码了多少?模型又能对其推理到什么程度?