Currently, the development of Foreign Accent Conversion (FAC) models utilizes deep neural network architectures, as well as ensembles of neural networks for speech recognition and speech generation. The use of these models is limited by architectural features, which does not allow flexible changes in the timbre of the generated speech and requires the accumulation of context, leading to increased delays in generation and makes these systems unsuitable for use in real-time multi-user communication scenarios. We have developed the non-autoregressive model for real-time accent conversion with voice cloning. The model generates native-sounding L1 speech with minimal latency based on input L2 accented speech. The model consists of interconnected modules for extracting accent, gender, and speaker embeddings, converting speech, generating spectrograms, and decoding the resulting spectrogram into an audio signal. The model has the ability to save, clone and change the timbre, gender and accent of the speaker's voice in real time. The results of the objective assessment show that the model improves speech quality, leading to enhanced recognition performance in existing ASR systems. The results of subjective tests show that the proposed accent and gender encoder improves the generation quality. The developed model demonstrates high-quality low-latency accent conversion, voice cloning, and speech enhancement capabilities, making it suitable for real-time multi-user communication scenarios.
翻译:目前,外语口音转换(FAC)模型的开发主要利用深度神经网络架构,以及用于语音识别和语音生成的神经网络集成。这些模型的使用受到架构特性的限制,无法灵活改变生成语音的音色,并且需要积累上下文,导致生成延迟增加,使得这些系统不适合用于实时多用户通信场景。我们开发了一种具有语音克隆功能的非自回归实时口音转换模型。该模型基于输入的L2口音语音,以最小延迟生成类似母语者的L1语音。该模型由相互关联的模块组成,用于提取口音、性别和说话人嵌入,转换语音,生成频谱图,并将生成的频谱图解码为音频信号。该模型能够实时保存、克隆并改变说话人语音的音色、性别和口音。客观评估结果表明,该模型提高了语音质量,从而提升了现有ASR系统的识别性能。主观测试结果表明,所提出的口音和性别编码器提高了生成质量。所开发的模型展示了高质量、低延迟的口音转换、语音克隆和语音增强能力,使其适用于实时多用户通信场景。