Humans effortlessly infer the 3D shape of objects. What computations underlie this ability? Although various computational models have been proposed, none of them capture the human ability to match object shape across viewpoints. Here, we ask whether and how this gap might be closed. We begin with a relatively novel class of computational models, 3D neural fields, which encapsulate the basic principles of classic analysis-by-synthesis in a deep neural network (DNN). First, we find that a 3D Light Field Network (3D-LFN) supports 3D matching judgments well aligned to humans for within-category comparisons, adversarially-defined comparisons that accentuate the 3D failure cases of standard DNN models, and adversarially-defined comparisons for algorithmically generated shapes with no category structure. We then investigate the source of the 3D-LFN's ability to achieve human-aligned performance through a series of computational experiments. Exposure to multiple viewpoints of objects during training and a multi-view learning objective are the primary factors behind model-human alignment; even conventional DNN architectures come much closer to human behavior when trained with multi-view objectives. Finally, we find that while the models trained with multi-view learning objectives are able to partially generalize to new object categories, they fall short of human alignment. This work provides a foundation for understanding human shape inferences within neurally mappable computational architectures and highlights important questions for future work.
翻译:人类能够毫不费力地推断物体的三维形状。何种计算机制支撑了这一能力?尽管已有多种计算模型被提出,但尚无模型能捕捉人类跨视角匹配物体形状的能力。本文探讨这一差距是否以及如何能被弥合。我们首先关注一类相对新颖的计算模型——三维神经场,它在深度神经网络中封装了经典分析-综合方法的基本原则。研究发现,三维光场网络在以下任务中支持与人类高度一致的三维匹配判断:类别内比较、突出标准DNN模型三维失败案例的对抗性定义比较,以及针对无类别结构的算法生成形状的对抗性比较。通过一系列计算实验,我们进一步探究了三维光场网络实现类人性能的根源:训练过程中对物体多视角的暴露以及多视角学习目标是模型与人类对齐的主要因素;即便是传统DNN架构,当采用多视角目标训练时,其行为也更接近人类。最后发现,尽管经多视角学习目标训练的模型能部分泛化至新物体类别,但仍未能达到人类对齐水平。本研究为在神经可映射计算架构中理解人类形状推理奠定了基础,并指出了未来研究的关键问题。