Recently Whisper has approached human-level robustness and accuracy in English speech recognition, while in minor language and mixed language speech recognition, there remains a compelling need for further improvement. In this work,we present the impressive results of Whisper-MCE, our fine-tuned Whisper, which was trainedusing our self-collected dataset, Mixed Cantoneseand English (MCE) audio dataset. Whisper-MCE achieved an impressive Mix Error Rate (MER) of 14.28%, which is 35.13% lower than the original model. It also achieved 12.61% Character Error Rate (CER) in Common voice zh-HK, positioning it as state-of-the-art. However, MER and CER pose challenges when it comes to evaluating its effectiveness in mixed-language and minor language contexts. We proposed a novel evaluation metric called FAL, which assesses an Automatic Speech Recognition (ASR) system based on fidelity to the original audio, accuracy, and latency. Whisper-MCE outperformed other models in this evaluation metric, achieving a score of 90.91 FAL, further highlighting its exceptional performance. The MCE dataset and code can be found at https://github.com/Shelton1013/Whisper MCE.
翻译:摘要:近期,Whisper在英语语音识别中已接近人类级别的鲁棒性与准确性,但在小语种及混合语言语音识别领域仍存在显著提升空间。本研究展示了经微调的Whisper-MCE模型(基于自采集的混合粤语与英语音频数据集MCE训练)的优异性能。Whisper-MCE实现了14.28%的混合错误率(MER),较原始模型降低35.13%,并在Common Voice zh-HK数据集上取得12.61%的字错误率(CER),达到当前最佳水平。然而,在混合语言与小语种场景下,MER与CER指标存在评估有效性不足的问题。为此,我们提出新型评估指标FAL,从音频保真度、准确性与延迟三个维度评估自动语音识别系统。Whisper-MCE在该指标下以90.91分超越其他模型,进一步凸显其卓越性能。MCE数据集与相关代码可于https://github.com/Shelton1013/Whisper MCE获取。