Standard transformer architectures learn fixed slow-weight representations during training and lack mechanisms for rapid adaptation within an episode. In contrast, biological neural systems address this through fast synaptic updates that form transient associative memories during inference, a property known as Hebbian plasticity. In this paper, we conduct an empirical study of Hebbian Fast-Weight (HFW) modules integrated into multiple transformer backbones, including ViT-Small, DeiT-Small, and Swin-Tiny. We evaluate six model variants: ViT, DeiT, Swin, ViT-Hebbian, DeiT-Hebbian, and Swin-Hebbian on 5-way 1-shot and 5-way 5-shot classification tasks using the Omniglot benchmark under a Prototypical Network meta-learning framework. We propose a single module placement strategy for Swin-Tiny in which one HFW module is applied to the final stage feature map after all hierarchical stages have completed. This design avoids the training instability caused by placing separate Hebbian modules at each stage and achieves the highest test accuracy across all six models (96.2\% at 1-shot; 99.2\% at 5-shot), outperforming its non-Hebbian baseline by $+0.3$ percentage points at 1-shot. We analyze the interaction between Swin's shifted window inductive bias and episode-level Hebbian binding, discuss why per-block placement fails for ViT and DeiT variants in a low-data regime, and situate the results within the wider literature on fast and slow-weight meta-learning.
翻译:标准Transformer架构在训练期间学习固定的慢权重表示,并缺乏在片段内快速适应的机制。与之相反,生物神经系统通过快速突触更新来形成瞬态关联记忆(一种称为赫布可塑性的特性)来解决这一问题。本文对集成到多个Transformer主干(包括ViT-Small、DeiT-Small和Swin-Tiny)中的赫布型快速权重模块进行了实证研究。我们评估了六种模型变体:ViT、DeiT、Swin、ViT-Hebbian、DeiT-Hebbian和Swin-Hebbian,在基于原型网络元学习框架的Omniglot基准测试上执行5-way 1-shot和5-way 5-shot分类任务。我们针对Swin-Tiny提出了一种单模块放置策略:在所有层级阶段完成后,仅将一个HFW模块应用于最终阶段的特征图。该设计避免了在每个阶段单独放置赫布模块导致的训练不稳定性,并在六种模型中取得了最高测试准确率(1-shot为96.2%;5-shot为99.2%),在1-shot任务上比其非赫布基线高出+0.3个百分点。我们分析了Swin的移位窗口归纳偏置与片段级赫布绑定的相互作用,讨论了在低数据场景下为何逐块放置对ViT和DeiT变体失败,并将这些结果置于关于快慢权重元学习的更广泛文献中。