Pre-trained language models, despite their rapid advancements powered by scale, still fall short of robust commonsense capabilities. And yet, scale appears to be the winning recipe; after all, the largest models seem to have acquired the largest amount of commonsense capabilities. Or is it? In this paper, we investigate the possibility of a seemingly impossible match: can smaller language models with dismal commonsense capabilities (i.e., GPT-2), ever win over models that are orders of magnitude larger and better (i.e., GPT-3), if the smaller models are powered with novel commonsense distillation algorithms? The key intellectual question we ask here is whether it is possible, if at all, to design a learning algorithm that does not benefit from scale, yet leads to a competitive level of commonsense acquisition. In this work, we study the generative models of commonsense knowledge, focusing on the task of generating generics, statements of commonsense facts about everyday concepts, e.g., birds can fly. We introduce a novel commonsense distillation framework, I2D2, that loosely follows the Symbolic Knowledge Distillation of West et al. but breaks the dependence on the extreme-scale models as the teacher model by two innovations: (1) the novel adaptation of NeuroLogic Decoding to enhance the generation quality of the weak, off-the-shelf language models, and (2) self-imitation learning to iteratively learn from the model's own enhanced commonsense acquisition capabilities. Empirical results suggest that scale is not the only way, as novel algorithms can be a promising alternative. Moreover, our study leads to a new corpus of generics, Gen-A-Tomic, that is of the largest and highest quality available to date.
翻译:预训练语言模型虽在规模驱动下快速进步,但仍缺乏稳健的常识推理能力。然而,规模似乎一直是制胜法宝——毕竟,最大规模的模型似乎掌握了最丰富的常识能力。但事实果真如此吗?本文探究一种看似不可能的对抗:当小型模型(如GPT-2)配备新型常识蒸馏算法时,其贫弱的常识能力能否击败规模大数个量级且性能更优的模型(如GPT-3)?我们提出的核心学术问题是:是否可能设计出一种不依赖模型规模却能实现竞争性常识习得水平的学习算法?本研究聚焦于常识知识生成模型,以生成“泛化陈述”(描述日常概念的常识事实陈述,例如“鸟会飞”)为任务。我们提出新型常识蒸馏框架I2D2,该框架虽大体遵循West等人的符号知识蒸馏范式,但通过两项创新打破了对超大规模教师模型的依赖:(1)新颖地适配NeuroLogic解码以增强弱预训练模型的生成质量;(2)采用自我模仿学习,使模型从自身增强的常识习得能力中迭代学习。实验结果表明,算法创新可成为规模扩展之外的有前景替代方案。此外,本研究还构建了迄今规模最大、质量最高的泛化陈述语料库Gen-A-Tomic。