Recent vision backbones, such as Transformer families and state-space models like Mamba, have achieved remarkable progress on image recognition. Despite their empirical success, these architectures remain far from the computational principles of the human brain, often demanding enormous amounts of training data while offering limited interpretability. We propose the Vision Hopfield Memory Network (V-HMN), a brain-inspired vision backbone that integrates hierarchical memory mechanisms across layers with iterative refinement updates. Specifically, V-HMN incorporates local Hopfield modules that provide associative memory dynamics at the image patch level, global Hopfield modules that function as episodic memory for contextual modulation, and a predictive-coding-inspired refinement rule for iterative error correction. By organizing these memory-based modules hierarchically, V-HMN captures both local and global dynamics in a unified framework. Memory retrieval exposes the relationship between inputs and stored patterns, providing a prototype-based form of interpretability through explicit memory retrieval, while the reuse of stored patterns improves data efficiency. This brain-inspired design therefore enhances data efficiency and provides a prototype-based form of interpretability compared to existing self-attention- or state-space-based approaches. We conducted extensive experiments on public image classification benchmarks. V-HMN achieves strong performance on small- and medium-scale benchmarks, and remains competitive with widely adopted backbone architectures on ImageNet despite minimal architectural tuning, while offering improved data efficiency and a prototype-based form of interpretability. These findings highlight the potential of V-HMN as a memory-centric alternative to standard vision backbones, thereby bridging brain-inspired computation with modern machine learning.


翻译:近期视觉骨干网络,如Transformer系列和基于状态空间模型的曼巴(Mamba),在图像识别领域取得了显著进展。尽管取得了经验上的成功,但这些架构仍远未达到人脑的计算原理,往往需要大量训练数据,且可解释性有限。我们提出了视觉霍普菲尔德记忆网络(V-HMN),这是一种受大脑启发的视觉骨干网络,它在各层之间集成了分层记忆机制与迭代精化更新。具体而言,V-HMN整合了局部霍普菲尔德模块(在图像块层面提供联想记忆动态)、全局霍普菲尔德模块(作为情境调制的情景记忆)以及基于预测编码的精化规则(用于迭代误差修正)。通过将这些基于记忆的模块分层组织,V-HMN在统一框架中捕获了局部和全局动态。记忆检索揭示了输入与存储模式之间的关系,通过显式记忆检索提供了一种基于原型的可解释性形式,而存储模式的重复利用则提高了数据效率。因此,与现有基于自注意力或状态空间的方法相比,这种受大脑启发的设计增强了数据效率,并提供了一种基于原型的可解释性形式。我们在公共图像分类基准上进行了大量实验。V-HMN在中小规模基准测试上取得了强劲性能,在ImageNet上尽管架构微调极少,仍能与广泛采用的骨干网络保持竞争力,同时提供了改进的数据效率和基于原型的可解释性。这些发现凸显了V-HMN作为标准视觉骨干网络的一种以记忆为中心的替代方案的潜力,从而弥合了受大脑启发的计算与现代机器学习之间的差距。

0
下载
关闭预览

相关内容

基于深度神经网络的高效视觉识别研究进展与新方向
专知会员服务
40+阅读 · 2021年8月31日
专知会员服务
11+阅读 · 2021年2月4日
【ICLR2020-】基于记忆的图网络,MEMORY-BASED GRAPH NETWORKS
专知会员服务
110+阅读 · 2020年2月22日
基于关系网络的视觉建模:有望替代卷积神经网络
微软研究院AI头条
10+阅读 · 2019年7月12日
【学界】基于条件深度卷积生成对抗网络的图像识别方法
GAN生成式对抗网络
16+阅读 · 2018年7月26日
基于注意力机制的图卷积网络
科技创新与创业
74+阅读 · 2017年11月8日
国家自然科学基金
6+阅读 · 2017年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
Arxiv
18+阅读 · 2019年3月28日
VIP会员
最新内容
一种考虑资源机动性的武器目标分配混合算法
专知会员服务
3+阅读 · 8月8日
《多域冲突比较支持模型》60页
专知会员服务
9+阅读 · 8月7日
面向2027年及未来的海军情报改革
专知会员服务
6+阅读 · 8月5日
相关基金
国家自然科学基金
6+阅读 · 2017年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员