English is the most widely spoken language in the world, used daily by millions of people as a first or second language in many different contexts. As a result, there are many varieties of English. Although the great many advances in English automatic speech recognition (ASR) over the past decades, results are usually reported based on test datasets which fail to represent the diversity of English as spoken today around the globe. We present the first release of The Edinburgh International Accents of English Corpus (EdAcc). This dataset attempts to better represent the wide diversity of English, encompassing almost 40 hours of dyadic video call conversations between friends. Unlike other datasets, EdAcc includes a wide range of first and second-language varieties of English and a linguistic background profile of each speaker. Results on latest public, and commercial models show that EdAcc highlights shortcomings of current English ASR models. The best performing model, trained on 680 thousand hours of transcribed data, obtains an average of 19.7% word error rate (WER) -- in contrast to the 2.7% WER obtained when evaluated on US English clean read speech. Across all models, we observe a drop in performance on Indian, Jamaican, and Nigerian English speakers. Recordings, linguistic backgrounds, data statement, and evaluation scripts are released on our website (https://groups.inf.ed.ac.uk/edacc/) under CC-BY-SA license.
翻译:英语是世界上使用最广泛的语言,每天有数百万人将其作为第一或第二语言在多种不同语境中使用,由此产生了众多英语变体。尽管过去几十年英语自动语音识别(ASR)取得了巨大进展,但相关结果通常基于未能反映当今全球英语多样性的测试数据集报告。我们发布了爱丁堡国际英语口音语料库(EdAcc)的首个版本。该数据集尝试更好地呈现英语的广泛多样性,包含近40小时的朋友间双人视频通话对话。与其他数据集不同,EdAcc涵盖了英语的多种第一语言和第二语言变体,并提供了每位说话者的语言背景档案。基于最新公开及商业模型的测试结果表明,EdAcc凸显了当前英语ASR模型的不足。其中表现最佳的模型在68万小时转录数据上训练,获得了19.7%的平均词错误率(WER)——而该模型在美国英语清晰朗读语音上的WER仅为2.7%。在所有模型中,我们观察到对印度、牙买加及尼日利亚英语使用者的性能下降。录音文件、语言背景信息、数据声明及评估脚本已以CC-BY-SA许可发布于我们的网站(https://groups.inf.ed.ac.uk/edacc/)。