We study speech enhancement using deep learning (DL) for virtual meetings on cellular devices, where transmitted speech has background noise and transmission loss that affects speech quality. Since the Deep Noise Suppression (DNS) Challenge dataset does not contain practical disturbance, we collect a transmitted DNS (t-DNS) dataset using Zoom Meetings over T-Mobile network. We select two baseline models: Demucs and FullSubNet. The Demucs is an end-to-end model that takes time-domain inputs and outputs time-domain denoised speech, and the FullSubNet takes time-frequency-domain inputs and outputs the energy ratio of the target speech in the inputs. The goal of this project is to enhance the speech transmitted over the cellular networks using deep learning models.
翻译:我们研究利用深度学习(DL)对蜂窝设备上虚拟会议中的语音进行增强,其中传输的语音包含背景噪声和影响语音质量的传输损失。由于深度噪声抑制(DNS)挑战数据集不包含实际干扰,我们使用T-Mobile网络上的Zoom会议收集了一个传输型DNS(t-DNS)数据集。我们选取了两个基线模型:Demucs和FullSubNet。Demucs是一种端到端模型,接收时域输入并输出时域去噪语音;FullSubNet则接收时频域输入,并输出输入中目标语音的能量比。本项目的目标是通过深度学习模型增强在蜂窝网络上传输的语音。