The Abstraction and Reasoning Corpus (ARC) (Chollet, 2019) and its most recent language-complete instantiation (LARC) has been postulated as an important step towards general AI. Yet, even state-of-the-art machine learning models struggle to achieve meaningful performance on these problems, falling behind non-learning based approaches. We argue that solving these tasks requires extreme generalization that can only be achieved by proper accounting for core knowledge priors. As a step towards this goal, we focus on geometry priors and introduce LatFormer, a model that incorporates lattice symmetry priors in attention masks. We show that, for any transformation of the hypercubic lattice, there exists a binary attention mask that implements that group action. Hence, our study motivates a modification to the standard attention mechanism, where attention weights are scaled using soft masks generated by a convolutional network. Experiments on synthetic geometric reasoning show that LatFormer requires 2 orders of magnitude fewer data than standard attention and transformers. Moreover, our results on ARC and LARC tasks that incorporate geometric priors provide preliminary evidence that these complex datasets do not lie out of the reach of deep learning models.
翻译:抽象与推理语料库(ARC)(Chollet, 2019)及其最新的语言完备化实例(LARC)被视为迈向通用人工智能的重要一步。然而,即便最先进的机器学习模型在这些问题上仍难以取得有意义的性能表现,落后于非学习方法。我们认为,解决这些任务需要极致的泛化能力,这只能通过恰当整合核心知识先验来实现。为实现这一目标,我们聚焦于几何先验,并提出了LatFormer模型,该模型将格点对称先验融入注意力掩码中。我们证明,对于超立方晶格的任意变换,均存在一个实现该群作用的二元注意力掩码。因此,本研究推动了对标准注意力机制的改进:通过卷积网络生成的软掩码对注意力权重进行缩放。在合成几何推理任务上的实验表明,LatFormer所需数据量比标准注意力和Transformer少两个数量级。此外,我们在包含几何先验的ARC和LARC任务上的结果为如下初步证据提供了支持:这些复杂数据集并未超出深度学习模型的能力范围。