Though modern neural networks have achieved impressive performance in both vision and language tasks, we know little about the functions that they implement. One possibility is that neural networks implicitly break down complex tasks into subroutines, implement modular solutions to these subroutines, and compose them into an overall solution to a task - a property we term structural compositionality. Another possibility is that they may simply learn to match new inputs to learned templates, eliding task decomposition entirely. Here, we leverage model pruning techniques to investigate this question in both vision and language across a variety of architectures, tasks, and pretraining regimens. Our results demonstrate that models often implement solutions to subroutines via modular subnetworks, which can be ablated while maintaining the functionality of other subnetworks. This suggests that neural networks may be able to learn compositionality, obviating the need for specialized symbolic mechanisms.
翻译:尽管现代神经网络在视觉和语言任务中均取得了显著性能,但我们对其实现的具体功能仍知之甚少。一种可能性是神经网络会隐式地将复杂任务分解为子程序,为这些子程序实现模块化解决方案,然后将它们组合成任务的整体解决方案——我们将此特性称为"结构组合性"。另一种可能性是神经网络可能仅通过学习将新输入与已学习的模板进行匹配,完全规避了任务分解过程。本研究利用模型剪枝技术,在多种架构、任务和预训练方案下,从视觉与语言两个领域对该问题展开探究。结果表明,模型通常通过模块化子网络实现子程序的解决方案——可在保持其他子网络功能的前提下对这些模块化子网络进行消融实验。这揭示出神经网络可能具备学习组合性的能力,从而无需借助专门的符号化机制。