Abstraction is a desirable capability for deep learning models, which means to induce abstract concepts from concrete instances and flexibly apply them beyond the learning context. At the same time, there is a lack of clear understanding about both the presence and further characteristics of this capability in deep learning models. In this paper, we introduce a systematic probing framework to explore the abstraction capability of deep learning models from a transferability perspective. A set of controlled experiments are conducted based on this framework, providing strong evidence that two probed pre-trained language models (PLMs), T5 and GPT2, have the abstraction capability. We also conduct in-depth analysis, thus shedding further light: (1) the whole training phase exhibits a "memorize-then-abstract" two-stage process; (2) the learned abstract concepts are gathered in a few middle-layer attention heads, rather than being evenly distributed throughout the model; (3) the probed abstraction capabilities exhibit robustness against concept mutations, and are more robust to low-level/source-side mutations than high-level/target-side ones; (4) generic pre-training is critical to the emergence of abstraction capability, and PLMs exhibit better abstraction with larger model sizes and data scales.
翻译:抽象是深度学习模型期望具备的能力,即从具体实例中归纳抽象概念并能在学习情境之外灵活应用。然而,关于深度学习模型是否具备该能力及其具体特征,目前仍缺乏清晰认知。本文提出了一个从可迁移性视角系统探测深度学习模型抽象能力的框架。基于该框架开展的一系列受控实验,为T5和GPT2两种预训练语言模型(PLMs)具备抽象能力提供了有力证据。通过深入分析,我们进一步揭示:(1)完整训练阶段呈现“先记忆后抽象”的两阶段过程;(2)习得的抽象概念集中于少数中间层注意力头,而非均匀分布于整个模型;(3)探测到的抽象能力对概念变异具有鲁棒性,且对低层级/源端变异的鲁棒性优于高层级/目标端变异;(4)通用预训练是抽象能力涌现的关键,更大模型规模与数据规模能提升PLMs的抽象能力表现。