Inspired by humans' exceptional ability to master arithmetic and generalize to new problems, we present a new dataset, Handwritten arithmetic with INTegers (HINT), to examine machines' capability of learning generalizable concepts at three levels: perception, syntax, and semantics. In HINT, machines are tasked with learning how concepts are perceived from raw signals such as images (i.e., perception), how multiple concepts are structurally combined to form a valid expression (i.e., syntax), and how concepts are realized to afford various reasoning tasks (i.e., semantics), all in a weakly supervised manner. Focusing on systematic generalization, we carefully design a five-fold test set to evaluate both the interpolation and the extrapolation of learned concepts w.r.t. the three levels. Further, we design a few-shot learning split to determine whether or not models can rapidly learn new concepts and generalize them to more complex scenarios. To comprehend existing models' limitations, we undertake extensive experiments with various sequence-to-sequence models, including RNNs, Transformers, and GPT-3 (with the chain of thought prompting). The results indicate that current models struggle to extrapolate to long-range syntactic dependency and semantics. Models exhibit a considerable gap toward human-level generalization when evaluated with new concepts in a few-shot setting. Moreover, we discover that it is infeasible to solve HINT by merely scaling up the dataset and the model size; this strategy contributes little to the extrapolation of syntax and semantics. Finally, in zero-shot GPT-3 experiments, the chain of thought prompting exhibits impressive results and significantly boosts the test accuracy. We believe the HINT dataset and the experimental findings are of great interest to the learning community on systematic generalization.
翻译:受人类在算术掌握及新问题泛化方面卓越能力的启发,我们提出了一个新数据集——手写整数算术(HINT),旨在从三个层次检验机器对可泛化概念的学习能力:感知层(如何从图像等原始信号中感知概念)、语法层(如何将多个概念结构组合为有效表达式)以及语义层(如何实现概念以完成各类推理任务),且所有学习均以弱监督方式进行。聚焦于系统泛化,我们精心设计了五折测试集,以评估所学概念在三个层次上的内插与外推能力。进一步地,我们设置了少样本学习划分,用以判断模型能否快速学习新概念并将其泛化至更复杂场景。为理解现有模型的局限性,我们使用包括RNN、Transformer及GPT-3(结合思维链提示)在内的多种序列到序列模型开展了大量实验。结果表明:当前模型难以外推至长程句法依赖与语义关系;在少样本场景下评估新概念时,模型与人类水平的泛化能力存在显著差距。此外,我们发现仅通过扩大数据集与模型规模无法解决HINT问题——该策略对句法与语义的外推贡献甚微。最后,在零样本GPT-3实验中,思维链提示展现出令人瞩目的效果,显著提升了测试准确率。我们相信,HINT数据集及其实验发现对系统泛化领域的学习社群具有重要参考价值。