We introduce O-1, a new self-training objective to reduce training bias and unify training and evaluation metrics for speech recognition. O-1 is a faster variant of Expected Minimum Bayes Risk (EMBR), that boosts the oracle hypothesis and can accommodate both supervised and unsupervised data. We demonstrate the effectiveness of our approach in terms of recognition on publicly available SpeechStew datasets and a large-scale, in-house data set. On Speechstew, the O-1 objective closes the gap between the actual and oracle performance by 80\% relative compared to EMBR which bridges the gap by 43\% relative. O-1 achieves 13\% to 25\% relative improvement over EMBR on the various datasets that SpeechStew comprises of, and a 12\% relative gap reduction with respect to the oracle WER over EMBR training on the in-house dataset. Overall, O-1 results in a 9\% relative improvement in WER over EMBR, thereby speaking to the scalability of the proposed objective for large-scale datasets.
翻译:我们提出O-1——一种新型自训练目标函数,旨在降低语音识别训练偏差并统一训练与评估指标。O-1是期望最小贝叶斯风险(EMBR)的快速变体,通过增强oracle假设同时支持监督与非监督数据。我们通过公开的SpeechStew数据集及大规模内部数据集验证了该方法的识别有效性。在SpeechStew上,O-1目标函数将实际性能与oracle性能之间的差距相对缩小80%,而EMBR仅相对缩小43%。在SpeechStew包含的多个数据集中,O-1较EMBR取得13%至25%的相对改进,并在内部数据集上相对于EMBR训练将oracle词错误率(WER)差距相对缩小12%。总体而言,O-1在WER上较EMBR实现9%的相对提升,充分表明该目标函数在大规模数据集上的可扩展性。