This paper describes the ESPnet Unsupervised ASR Open-source Toolkit (EURO), an end-to-end open-source toolkit for unsupervised automatic speech recognition (UASR). EURO adopts the state-of-the-art UASR learning method introduced by the Wav2vec-U, originally implemented at FAIRSEQ, which leverages self-supervised speech representations and adversarial training. In addition to wav2vec2, EURO extends the functionality and promotes reproducibility for UASR tasks by integrating S3PRL and k2, resulting in flexible frontends from 27 self-supervised models and various graph-based decoding strategies. EURO is implemented in ESPnet and follows its unified pipeline to provide UASR recipes with a complete setup. This improves the pipeline's efficiency and allows EURO to be easily applied to existing datasets in ESPnet. Extensive experiments on three mainstream self-supervised models demonstrate the toolkit's effectiveness and achieve state-of-the-art UASR performance on TIMIT and LibriSpeech datasets. EURO will be publicly available at https://github.com/espnet/espnet, aiming to promote this exciting and emerging research area based on UASR through open-source activity.
翻译:本文介绍了ESPnet无监督自动语音识别开源工具包(EURO),这是一个用于无监督自动语音识别(UASR)的端到端开源工具包。EURO采用了Wav2vec-U提出的最先进UASR学习方法(最初在FAIRSEQ中实现),该方法利用了自监督语音表示和对抗性训练。除wav2vec2外,EURO通过集成S3PRL和k2扩展了UASR任务的功能并提升了可复现性,实现了来自27个自监督模型的灵活前端以及多种基于图模型的解码策略。EURO基于ESPnet实现,并遵循其统一流水线以提供完整配置的UASR配方。这提升了流水线的效率,使EURO能够轻松应用于ESPnet中的现有数据集。在三种主流自监督模型上的大量实验证明了该工具包的有效性,并在TIMIT和LibriSpeech数据集上实现了最先进的UASR性能。EURO将在https://github.com/espnet/espnet公开发布,旨在通过开源活动推动这一基于UASR的激动人心的新兴研究领域。