Compositional generalization, the ability to predict complex meanings from training on simpler sentences, poses challenges for powerful pretrained seq2seq models. In this paper, we show that data augmentation methods that sample MRs and backtranslate them can be effective for compositional generalization, but only if we sample from the right distribution. Remarkably, sampling from a uniform distribution performs almost as well as sampling from the test distribution, and greatly outperforms earlier methods that sampled from the training distribution. We further conduct experiments to investigate the reason why this happens and where the benefit of such data augmentation methods come from.
翻译:组合泛化,即从简单句训练中预测复杂含义的能力,对强大的预训练序列到序列模型构成挑战。本文证明,对意义表示进行采样并回译的数据增强方法可有效提升组合泛化能力,但前提是采样需来自正确的分布。值得注意的是,从均匀分布中采样的效果几乎与从测试分布中采样相当,且大幅优于以往从训练分布中采样的方法。我们进一步通过实验探究这一现象背后的原因,以及此类数据增强方法益处的来源。