Encoding static images into spike trains is a fundamental step for enabling Spiking Neural Networks (SNNs) to process visual information. However, widely used methods such as rate coding, Poisson encoding, and time-to-first-spike (TTFS) often neglect spatial correlations and produce temporally inconsistent spike patterns, limiting both efficiency and interpretability. In this work, we propose a novel cluster-based encoding framework that explicitly preserves semantic structure across both spatial and temporal domains. The method first introduces a 2D spatial clustering mechanism, which leverages connected component analysis and local density estimation to identify salient foreground regions. Building upon this, we extend the approach to a 3D spatio-temporal (ST3D) encoding scheme that incorporates temporal neighborhood information, generating spike trains with enhanced temporal coherence. Experiments on the N-MNIST dataset demonstrate that the proposed ST3D encoder achieves 98.17% classification accuracy using a simple single-layer SNN, outperforming conventional TTFS encoding (97.58%). Notably, this performance is achieved with significantly fewer spikes (3800 vs. 5000 per sample), highlighting improved efficiency without sacrificing accuracy. These results indicate that the proposed method provides an interpretable, structure-aware, and computationally efficient encoding strategy, offering strong potential for neuromorphic computing applications.
翻译:将静态图像编码为脉冲序列是使脉冲神经网络(SNN)处理视觉信息的基础步骤。然而,速率编码、泊松编码和首次脉冲时间编码(TTFS)等广泛采用的方法常忽略空间相关性,并产生时间上不一致的脉冲模式,从而限制了编码效率与可解释性。本文提出一种新颖的基于聚类的编码框架,该框架在空间和时间两个维度上显式保留语义结构。该方法首先引入二维空间聚类机制,利用连通分量分析和局部密度估计来识别显著的感兴趣前景区域。在此基础上,我们将该方法扩展为三维时空(ST3D)编码方案,该方案融入时间邻域信息,生成具有更强时间一致性的脉冲序列。在N-MNIST数据集上的实验表明,所提出的ST3D编码器采用简单的单层SNN即可实现98.17%的分类准确率,优于传统TTFS编码方法(97.58%)。值得注意的是,该性能是在使用显著更少脉冲(每样本3800个对比5000个)的情况下实现的,凸显了效率提升而未牺牲精度。这些结果表明,所提方法提供了一种可解释、结构感知且计算高效的编码策略,为神经形态计算应用展现了巨大潜力。