Recent vision backbones, such as Transformer families and state-space models like Mamba, have achieved remarkable progress on image recognition. Despite their empirical success, these architectures remain far from the computational principles of the human brain, often demanding enormous amounts of training data while offering limited interpretability. We propose the Vision Hopfield Memory Network (V-HMN), a brain-inspired vision backbone that integrates hierarchical memory mechanisms across layers with iterative refinement updates. Specifically, V-HMN incorporates local Hopfield modules that provide associative memory dynamics at the image patch level, global Hopfield modules that function as episodic memory for contextual modulation, and a predictive-coding-inspired refinement rule for iterative error correction. By organizing these memory-based modules hierarchically, V-HMN captures both local and global dynamics in a unified framework. Memory retrieval exposes the relationship between inputs and stored patterns, providing a prototype-based form of interpretability through explicit memory retrieval, while the reuse of stored patterns improves data efficiency. This brain-inspired design therefore enhances data efficiency and provides a prototype-based form of interpretability compared to existing self-attention- or state-space-based approaches. We conducted extensive experiments on public image classification benchmarks. V-HMN achieves strong performance on small- and medium-scale benchmarks, and remains competitive with widely adopted backbone architectures on ImageNet despite minimal architectural tuning, while offering improved data efficiency and a prototype-based form of interpretability. These findings highlight the potential of V-HMN as a memory-centric alternative to standard vision backbones, thereby bridging brain-inspired computation with modern machine learning.
翻译:近期视觉骨干网络,如Transformer系列和基于状态空间模型的曼巴(Mamba),在图像识别领域取得了显著进展。尽管取得了经验上的成功,但这些架构仍远未达到人脑的计算原理,往往需要大量训练数据,且可解释性有限。我们提出了视觉霍普菲尔德记忆网络(V-HMN),这是一种受大脑启发的视觉骨干网络,它在各层之间集成了分层记忆机制与迭代精化更新。具体而言,V-HMN整合了局部霍普菲尔德模块(在图像块层面提供联想记忆动态)、全局霍普菲尔德模块(作为情境调制的情景记忆)以及基于预测编码的精化规则(用于迭代误差修正)。通过将这些基于记忆的模块分层组织,V-HMN在统一框架中捕获了局部和全局动态。记忆检索揭示了输入与存储模式之间的关系,通过显式记忆检索提供了一种基于原型的可解释性形式,而存储模式的重复利用则提高了数据效率。因此,与现有基于自注意力或状态空间的方法相比,这种受大脑启发的设计增强了数据效率,并提供了一种基于原型的可解释性形式。我们在公共图像分类基准上进行了大量实验。V-HMN在中小规模基准测试上取得了强劲性能,在ImageNet上尽管架构微调极少,仍能与广泛采用的骨干网络保持竞争力,同时提供了改进的数据效率和基于原型的可解释性。这些发现凸显了V-HMN作为标准视觉骨干网络的一种以记忆为中心的替代方案的潜力,从而弥合了受大脑启发的计算与现代机器学习之间的差距。