We present dual-attention neural biasing, an architecture designed to boost Wake Words (WW) recognition and improve inference time latency on speech recognition tasks. This architecture enables a dynamic switch for its runtime compute paths by exploiting WW spotting to select which branch of its attention networks to execute for an input audio frame. With this approach, we effectively improve WW spotting accuracy while saving runtime compute cost as defined by floating point operations (FLOPs). Using an in-house de-identified dataset, we demonstrate that the proposed dual-attention network can reduce the compute cost by $90\%$ for WW audio frames, with only $1\%$ increase in the number of parameters. This architecture improves WW F1 score by $16\%$ relative and improves generic rare word error rate by $3\%$ relative compared to the baselines.
翻译:我们提出双注意力神经偏向架构,旨在提升语音识别任务中唤醒词的识别准确率并降低推理时间延迟。该架构通过利用唤醒词检测技术,为输入音频帧动态切换注意力网络的执行分支,从而优化运行时计算路径。采用此方法,我们有效提升了唤醒词识别精度,同时以浮点运算数衡量的运行时计算成本得以显著降低。基于内部去标识化数据集的实验表明,所提出的双注意力网络可将唤醒词语音帧的计算成本降低90%,而参数量仅增加1%。与基线模型相比,该架构使唤醒词F1分数相对提升16%,罕见词错误率相对降低3%。