Voice communication between air traffic controllers (ATCos) and pilots is critical for ensuring safe and efficient air traffic control (ATC). This task requires high levels of awareness from ATCos and can be tedious and error-prone. Recent attempts have been made to integrate artificial intelligence (AI) into ATC in order to reduce the workload of ATCos. However, the development of data-driven AI systems for ATC demands large-scale annotated datasets, which are currently lacking in the field. This paper explores the lessons learned from the ATCO2 project, a project that aimed to develop a unique platform to collect and preprocess large amounts of ATC data from airspace in real time. Audio and surveillance data were collected from publicly accessible radio frequency channels with VHF receivers owned by a community of volunteers and later uploaded to Opensky Network servers, which can be considered an "unlimited source" of data. In addition, this paper reviews previous work from ATCO2 partners, including (i) robust automatic speech recognition, (ii) natural language processing, (iii) English language identification of ATC communications, and (iv) the integration of surveillance data such as ADS-B. We believe that the pipeline developed during the ATCO2 project, along with the open-sourcing of its data, will encourage research in the ATC field. A sample of the ATCO2 corpus is available on the following website: https://www.atco2.org/data, while the full corpus can be purchased through ELDA at http://catalog.elra.info/en-us/repository/browse/ELRA-S0484. We demonstrated that ATCO2 is an appropriate dataset to develop ASR engines when little or near to no ATC in-domain data is available. For instance, with the CNN-TDNNf kaldi model, we reached the performance of as low as 17.9% and 24.9% WER on public ATC datasets which is 6.6/7.6% better than "out-of-domain" but supervised CNN-TDNNf model.
翻译:空中交通管制员(ATCos)与飞行员之间的语音通信对于确保安全高效的空管(ATC)至关重要。该任务要求ATCos保持高度警觉,且可能繁琐且易出错。近年来,人们尝试将人工智能(AI)集成到空管中以减轻ATCos的工作负担。然而,开发基于数据的空管AI系统需要大规模带注释的数据集,目前该领域尚缺乏此类资源。本文探讨了ATCO2项目的经验总结,该项目旨在开发一个独特平台,用于实时收集和预处理大量空域中的空管数据。音频和监视数据通过志愿者社区拥有的甚高频接收机,从公开可用的射频频道收集,随后上传至Opensky Network服务器,该服务器可被视为数据的“无限来源”。此外,本文回顾了ATCO2合作伙伴的先前工作,包括:(i)鲁棒自动语音识别,(ii)自然语言处理,(iii)空管通信的英语语种识别,以及(iv)ADS-B等监视数据的集成。我们相信,ATCO2项目开发的流水线及其数据的开源,将推动空管领域的研究。ATCO2语料库样本可在以下网站获取:https://www.atco2.org/data,完整语料库可通过ELDA购买,网址为http://catalog.elra.info/en-us/repository/browse/ELRA-S0484。我们证明,在几乎没有空管领域内数据可用时,ATCO2是开发ASR引擎的合适数据集。例如,使用CNN-TDNNf kaldi模型,我们在公开空管数据集上实现了低至17.9%和24.9%的词错误率(WER),比“领域外”但有监督的CNN-TDNNf模型分别提升了6.6%和7.6%。