Real-life multilingual systems should be able to efficiently incorporate new languages as data distributions fed to the system evolve and shift over time. To do this, systems need to handle the issue of catastrophic forgetting, where the model performance drops for languages or tasks seen further in its past. In this paper, we study catastrophic forgetting, as well as methods to minimize this, in a massively multilingual continual learning framework involving up to 51 languages and covering both classification and sequence labeling tasks. We present LR ADJUST, a learning rate scheduling method that is simple, yet effective in preserving new information without strongly overwriting past knowledge. Furthermore, we show that this method is effective across multiple continual learning approaches. Finally, we provide further insights into the dynamics of catastrophic forgetting in this massively multilingual setup.
翻译:实际应用中的多语言系统需要能够随着输入系统的数据分布演变和变化,高效地纳入新语言。为此,系统必须应对灾难性遗忘问题——即模型对较早接触的语言或任务的性能下降。本文在一个包含多达51种语言、涵盖分类和序列标注任务的大规模多语言持续学习框架中,研究了灾难性遗忘及其缓解方法。我们提出了一种名为LR ADJUST的学习率调度方法,该方法简单且有效,能在不强烈覆盖已有知识的情况下保留新信息。此外,我们证明该方法在多种持续学习方案中均有效。最后,我们进一步深入分析了这种大规模多语言设置下灾难性遗忘的动态机制。