Voice assistants increasingly use on-device Automatic Speech Recognition (ASR) to ensure speed and privacy. However, due to resource constraints on the device, queries pertaining to complex information domains often require further processing by a search engine. For such applications, we propose a novel Transformer based model capable of rescoring and rewriting, by exploring full context of the N-best hypotheses in parallel. We also propose a new discriminative sequence training objective that can work well for both rescore and rewrite tasks. We show that our Rescore+Rewrite model outperforms the Rescore-only baseline, and achieves up to an average 8.6% relative Word Error Rate (WER) reduction over the ASR system by itself.
翻译:为确保响应速度与隐私保护,语音助手日益采用设备端自动语音识别技术。然而,由于设备资源受限,涉及复杂信息领域的查询通常需要搜索引擎进行进一步处理。针对此类应用,我们提出一种基于Transformer的新型模型,该模型能够通过并行探索N-Best假设的完整上下文,同时执行重打分与改写任务。我们还提出了一种新的判别式序列训练目标,该目标能同时适用于重打分与改写任务。实验表明,我们的"重打分+改写"模型优于仅进行重打分的基线系统,相比原始ASR系统平均实现了高达8.6%的相对词错误率降低。