Spiking neural networks (SNNs) have tremendous potential for energy-efficient neuromorphic chips due to their binary and event-driven architecture. SNNs have been primarily used in classification tasks, but limited exploration on image generation tasks. To fill the gap, we propose a Spiking-Diffusion model, which is based on the vector quantized discrete diffusion model. First, we develop a vector quantized variational autoencoder with SNNs (VQ-SVAE) to learn a discrete latent space for images. With VQ-SVAE, image features are encoded using both the spike firing rate and postsynaptic potential, and an adaptive spike generator is designed to restore embedding features in the form of spike trains. Next, we perform absorbing state diffusion in the discrete latent space and construct a diffusion image decoder with SNNs to denoise the image. Our work is the first to build the diffusion model entirely from SNN layers. Experimental results on MNIST, FMNIST, KMNIST, and Letters demonstrate that Spiking-Diffusion outperforms the existing SNN-based generation model. We achieve FIDs of 37.50, 91.98, 59.23 and 67.41 on the above datasets respectively, with reductions of 58.60\%, 18.75\%, 64.51\%, and 29.75\% in FIDs compared with the state-of-art work.
翻译:脉冲神经网络(SNN)因其二进制和事件驱动架构,在节能神经形态芯片方面具有巨大潜力。目前SNN主要应用于分类任务,但在图像生成任务中的探索十分有限。为填补这一空白,我们提出了一种基于矢量量化离散扩散模型的脉冲扩散模型。首先,我们开发了基于SNN的矢量量化变分自编码器(VQ-SVAE),用于学习图像的离散潜在空间。借助VQ-SVAE,图像特征通过脉冲发放率和突触后电位进行编码,并设计自适应脉冲生成器以脉冲序列形式恢复嵌入特征。其次,我们在离散潜在空间中执行吸收态扩散,并构建基于SNN的扩散图像解码器进行去噪。本工作首次完全使用SNN层构建扩散模型。在MNIST、FMNIST、KMNIST和Letters数据集上的实验结果表明,脉冲扩散模型优于现有基于SNN的生成模型。在上述数据集上,我们分别获得了37.50、91.98、59.23和67.41的FID分数,相较于当前最优方法,FID分数分别降低了58.60%、18.75%、64.51%和29.75%。