Representation forgetting refers to the drift of contextualized representations during continual training. Intuitively, the representation forgetting can influence the general knowledge stored in pre-trained language models (LMs), but the concrete effect is still unclear. In this paper, we study the effect of representation forgetting on the generality of pre-trained language models, i.e. the potential capability for tackling future downstream tasks. Specifically, we design three metrics, including overall generality destruction (GD), syntactic knowledge forgetting (SynF), and semantic knowledge forgetting (SemF), to measure the evolution of general knowledge in continual learning. With extensive experiments, we find that the generality is destructed in various pre-trained LMs, and syntactic and semantic knowledge is forgotten through continual learning. Based on our experiments and analysis, we further get two insights into alleviating general knowledge forgetting: 1) training on general linguistic tasks at first can mitigate general knowledge forgetting; 2) the hybrid continual learning method can mitigate the generality destruction and maintain more general knowledge compared with those only considering rehearsal or regularization.
翻译:表征遗忘指的是在持续训练过程中,上下文表征发生漂移的现象。直观上,表征遗忘可能影响预训练语言模型中存储的通用知识,但其具体影响尚不明确。本文研究了表征遗忘对预训练语言模型通用性(即处理未来下游任务的潜在能力)的影响。具体而言,我们设计了三个度量指标:通用性破坏(GD)、句法知识遗忘(SynF)和语义知识遗忘(SemF),用以衡量通用知识在持续学习中的演变过程。通过大量实验发现,各类预训练语言模型的通用性均受到破坏,句法和语义知识在持续学习中逐渐被遗忘。基于实验与分析,我们进一步提出缓解通用知识遗忘的两点启示:1) 优先训练通用语言任务可缓解通用知识遗忘;2) 相较于仅采用重放或正则化的方法,混合持续学习方法能减轻通用性破坏,并保留更多通用知识。