Referring expression segmentation aims to segment an object described by a language expression from an image. Despite the recent progress on this task, existing models tackling this task may not be able to fully capture semantics and visual representations of individual concepts, which limits their generalization capability, especially when handling novel compositions of learned concepts. In this work, through the lens of meta learning, we propose a Meta Compositional Referring Expression Segmentation (MCRES) framework to enhance model compositional generalization performance. Specifically, to handle various levels of novel compositions, our framework first uses training data to construct a virtual training set and multiple virtual testing sets, where data samples in each virtual testing set contain a level of novel compositions w.r.t. the virtual training set. Then, following a novel meta optimization scheme to optimize the model to obtain good testing performance on the virtual testing sets after training on the virtual training set, our framework can effectively drive the model to better capture semantics and visual representations of individual concepts, and thus obtain robust generalization performance even when handling novel compositions. Extensive experiments on three benchmark datasets demonstrate the effectiveness of our framework.
翻译:指代分割旨在从图像中分割出由语言表达式所描述的物体。尽管该任务近期取得了进展,但现有的处理方法可能无法充分捕捉个体概念的语义与视觉表征,这限制了它们的泛化能力,尤其是在处理已习得概念的新颖组合时。在本工作中,我们通过元学习的视角,提出了一种元组合式指代分割(MCRES)框架,以增强模型的组合泛化性能。具体而言,为应对不同级别的新颖组合,我们的框架首先利用训练数据构建一个虚拟训练集和多个虚拟测试集,其中每个虚拟测试集中的数据样本相对于虚拟训练集包含一定程度的新颖组合。随后,遵循一种新颖的元优化方案,在虚拟训练集上训练模型后,使其在虚拟测试集上获得良好的测试性能,我们的框架能够有效驱动模型更好地捕捉个体概念的语义与视觉表征,从而在处理新颖组合时也获得稳健的泛化性能。在三个基准数据集上的大量实验证明了我们框架的有效性。