In today's interconnected globe, moving abroad is more and more prevalent, whether it's for employment, refugee resettlement, or other causes. Language difficulties between natives and immigrants present a common issue on a daily basis, especially in medical domain. This can make it difficult for patients and doctors to communicate during anamnesis or in the emergency room, which compromises patient care. The goal of the HYKIST Project is to develop a speech translation system to support patient-doctor communication with ASR and MT. ASR systems have recently displayed astounding performance on particular tasks for which enough quantities of training data are available, such as LibriSpeech. Building a good model is still difficult due to a variety of speaking styles, acoustic and recording settings, and a lack of in-domain training data. In this thesis, we describe our efforts to construct ASR systems for a conversational telephone speech recognition task in the medical domain for Vietnamese language to assist emergency room contact between doctors and patients across linguistic barriers. In order to enhance the system's performance, we investigate various training schedules and data combining strategies. We also examine how best to make use of the little data that is available. The use of publicly accessible models like XLSR-53 is compared to the use of customized pre-trained models, and both supervised and unsupervised approaches are utilized using wav2vec 2.0 as architecture.
翻译:在当今互联互通的世界中,无论是出于就业、难民安置还是其他原因,移居国外越来越普遍。本地人与移民之间的语言障碍在日常交流中构成常见问题,尤其在医疗领域尤为突出。这可能导致患者在病史采集或急诊室环境中与医生沟通困难,从而影响患者护理质量。HYKIST项目旨在开发一套融合自动语音识别(ASR)与机器翻译(MT)的语音翻译系统,以支持医患沟通。近年来,在拥有充足训练数据的特定任务(如LibriSpeech)上,ASR系统展现出卓越性能。然而,由于说话风格、声学环境与录音条件的多样性,以及领域内训练数据的匮乏,构建优质模型仍面临挑战。本文描述了针对越南语医疗领域对话式电话语音识别任务构建ASR系统的相关努力,旨在跨越语言障碍协助急诊室医患沟通。为提升系统性能,我们探索了多种训练策略与数据组合方案,并研究了如何最优利用有限的可用数据。此外,我们对比了使用XLSR-53等公开模型与定制化预训练模型的效果,并基于wav2vec 2.0架构同时采用监督与无监督方法。