Natively trained spiking language models struggle to combine Transformer-like language quality, stable multi-domain pre-training, and high activation sparsity. We present SymbolicLight V1, a spike-gated dual-path language model that combines binary Leaky Integrate-and-Fire spike dynamics with a continuous residual stream. Its Dual-Path SparseTCAM module replaces dense self-attention with an exponential-decay aggregation path for long-range memory and a spike-gated local attention path for short-range precision, complemented by a dynamic context-conditioned decoding head and a bilingual tokenizer. A 194M-parameter SymbolicLight V1 model trained from scratch on a 3B-token Chinese-English corpus reaches held-out validation PPL 8.88-8.93 across four independent runs at >89% per-element activation sparsity. It trails GPT-2 201M by 7.7% in PPL while surpassing GPT-2 124M under the reported comparison. Component ablations at matched 0.5B-token training budgets show that the spike-gated local attention path is the largest contributor, and that replacing LIF dynamics with a deterministic top-k mask at matched sparsity causes a larger degradation, indicating that temporal integration rather than sparsity alone drives performance. We also report a 0.8B-parameter scale-up run trained on 48.8B tokens as evidence of optimization and sparsity preservation, not as a primary quality comparison. Current dense-hardware inference is slower than GPT-2, so neuromorphic deployment is presented as a future sparsity-driven opportunity rather than an achieved hardware speedup.
翻译:原生训练的脉冲语言模型难以兼顾类似Transformer的语言质量、稳定的多领域预训练以及高激活稀疏性。我们提出SymbolicLight V1,一种脉冲门控双路径语言模型,它将二元漏积分-放电脉冲动力学与连续残差流相结合。其双路径稀疏TCAM模块采用指数衰减聚合路径用于长程记忆、脉冲门控局部注意力路径用于短程精度,从而替代了密集自注意力,并辅以动态上下文条件解码头及双语分词器。一个在30亿词次中英文语料上从头训练的1.94亿参数SymbolicLight V1模型,在四次独立运行中,以超过89%的逐元素激活稀疏性取得了8.88-8.93的留出验证集困惑度。在报道的对比下,其困惑度虽比GPT-2 201M模型高7.7%,但超越了GPT-2 124M模型。在匹配的5亿词次训练预算下进行的组件消融实验表明,脉冲门控局部注意力路径贡献最大;而在匹配稀疏度下,用确定性top-k掩码替换LIF动力学会导致更大的性能下降,这表明是时序整合而非单纯的稀疏性驱动了性能提升。我们还报告了一个在488亿词次上训练的8亿参数规模扩展实验,作为优化与稀疏性保持的证据,而非主要的质量对比。当前密集硬件推理速度慢于GPT-2,因此神经形态部署作为未来稀疏性驱动的机遇被提出,而非已实现的硬件加速。