We improve on the popular conformer architecture by replacing the depthwise temporal convolutions with diagonal state space (DSS) models. DSS is a recently introduced variant of linear RNNs obtained by discretizing a linear dynamical system with a diagonal state transition matrix. DSS layers project the input sequence onto a space of orthogonal polynomials where the choice of basis functions, metric and support is controlled by the eigenvalues of the transition matrix. We compare neural transducers with either conformer or our proposed DSS-augmented transformer (DSSformer) encoders on three public corpora: Switchboard English conversational telephone speech 300 hours, Switchboard+Fisher 2000 hours, and a spoken archive of holocaust survivor testimonials called MALACH 176 hours. On Switchboard 300/2000 hours, we reach a single model performance of 8.9%/6.7% WER on the combined test set of the Hub5 2000 evaluation, respectively, and on MALACH we improve the WER by 7% relative over the previous best published result. In addition, we present empirical evidence suggesting that DSS layers learn damped Fourier basis functions where the attenuation coefficients are layer specific whereas the frequency coefficients converge to almost identical linearly-spaced values across all layers.
翻译:我们对流行的 Conformer 架构进行了改进,用对角状态空间(DSS)模型替换了深度可分离时间卷积。DSS 是最近引入的一种线性循环神经网络变体,通过对具有对角状态转移矩阵的线性动态系统进行离散化得到。DSS 层将输入序列投影到正交多项式空间上,其中基函数、度量和支撑集的选择由转移矩阵的特征值控制。我们在三个公开语料库上比较了使用 Conformer 编码器与我们所提出的 DSS 增强 Transformer(DSSformer)编码器的神经换能器:Switchboard 英语会话电话语音(300 小时)、Switchboard+Fisher(2000 小时)以及名为 MALACH(176 小时)的大屠杀幸存者证词汇录音档案。在 Switchboard 300/2000 小时数据集上,我们在 Hub5 2000 评估的联合测试集上分别达到了 8.9%/6.7% 的词错误率(WER)的单模型性能;在 MALACH 上,我们相比之前最佳已发表结果将 WER 相对降低了 7%。此外,我们提供的实证证据表明,DSS 层学习的是阻尼傅里叶基函数,其中衰减系数因层而异,而频率系数在所有层中收敛到几乎相同的等间距值。