Among the array of neural network architectures, the Vision Transformer (ViT) stands out as a prominent choice, acclaimed for its exceptional expressiveness and consistent high performance in various vision applications. Recently, the emerging Spiking ViT approach has endeavored to harness spiking neurons, paving the way for a more brain-inspired transformer architecture that thrives in ultra-low power operations on dedicated neuromorphic hardware. Nevertheless, this approach remains confined to spatial self-attention and doesn't fully unlock the potential of spiking neural networks. We introduce DISTA, a Denoising Spiking Transformer with Intrinsic Plasticity and SpatioTemporal Attention, designed to maximize the spatiotemporal computational prowess of spiking neurons, particularly for vision applications. DISTA explores two types of spatiotemporal attentions: intrinsic neuron-level attention and network-level attention with explicit memory. Additionally, DISTA incorporates an efficient nonlinear denoising mechanism to quell the noise inherent in computed spatiotemporal attention maps, thereby resulting in further performance gains. Our DISTA transformer undergoes joint training involving synaptic plasticity (i.e., weight tuning) and intrinsic plasticity (i.e., membrane time constant tuning) and delivers state-of-the-art performances across several static image and dynamic neuromorphic datasets. With only 6 time steps, DISTA achieves remarkable top-1 accuracy on CIFAR10 (96.26%) and CIFAR100 (79.15%), as well as 79.1% on CIFAR10-DVS using 10 time steps.
翻译:在众多神经网络架构中,Vision Transformer(ViT)因其在各种视觉应用中卓越的表现力和持续的高性能而成为突出选择。近期兴起的脉冲ViT方法尝试利用脉冲神经元,为构建更受大脑启发的Transformer架构铺平了道路,使其能在专用神经形态硬件上以超低功耗运行。然而,该方法仍局限于空间自注意力,未能完全释放脉冲神经网络的潜力。我们提出DISTA——一种具有内在可塑性和时空注意力的去噪脉冲Transformer,旨在最大化脉冲神经元在视觉应用中的时空计算能力。DISTA探索了两种时空注意力机制:内在神经元级注意力和带有显式记忆的网络级注意力。此外,DISTA集成了高效的非线性去噪机制,以抑制计算所得时空注意力图中的固有噪声,从而进一步提升性能。我们的DISTA Transformer通过联合训练突触可塑性(即权重调整)和内在可塑性(即膜时间常数调整),在多个静态图像和动态神经形态数据集上实现了最先进的性能。仅用6个时间步,DISTA在CIFAR10(96.26%)和CIFAR100(79.15%)上便达到了卓越的top-1准确率,在CIFAR10-DVS上用10个时间步实现了79.1%的准确率。