The degree to which neural networks can generalize to new combinations of familiar concepts, and the conditions under which they are able to do so, has long been an open question. In this work, we study the systematicity gap in visual question answering: the performance difference between reasoning on previously seen and unseen combinations of object attributes. To test, we introduce a novel diagnostic dataset, CLEVR-HOPE. We find that while increased quantity of training data does not reduce the systematicity gap, increased training data diversity of the attributes in the unseen combination does. In all, our experiments suggest that the more distinct attribute type combinations are seen during training, the more systematic we can expect the resulting model to be.
翻译:神经网络在多大程度上能够泛化到熟悉概念的新组合,以及它们在何种条件下能够做到这一点,长期以来一直是一个悬而未决的问题。本研究聚焦于视觉问答中的系统性差距:对先前见过与未见过的对象属性组合进行推理时的性能差异。为进行测试,我们提出了一个新颖的诊断数据集CLEVR-HOPE。我们发现,虽然增加训练数据量并不能缩小系统性差距,但提高未见组合中属性的训练数据多样性却能实现这一目标。总体而言,我们的实验表明,训练过程中观察到的不同属性类型组合越多,最终模型的系统性就越强。