This research is about the creation of personalized synthetic voices for head and neck cancer survivors. It is focused particularly on tongue cancer patients whose speech might exhibit severe articulation impairment. Our goal is to restore normal articulation in the synthesized speech, while maximally preserving the target speaker's individuality in terms of both the voice timbre and speaking style. This is formulated as a task of learning from noisy labels. We propose to augment the commonly used speech reconstruction loss with two additional terms. The first term constitutes a regularization loss that mitigates the impact of distorted articulation in the training speech. The second term is a consistency loss that encourages correct articulation in the generated speech. These additional loss terms are obtained from frame-level articulation scores of original and generated speech, which are derived using a separately trained phone classifier. Experimental results on a real case of tongue cancer patient confirm that the synthetic voice achieves comparable articulation quality to unimpaired natural speech, while effectively maintaining the target speaker's individuality. Audio samples are available at https://myspeechproject.github.io/ArticulationRepair/.
翻译:本研究旨在为头颈部癌症幸存者创建个性化合成声音,特别关注舌癌患者,其语音可能表现出严重的构音障碍。我们的目标是在合成语音中恢复正常的构音,同时最大限度地保留目标说话人在语音音色和说话风格方面的个性。这被形式化为一个从噪声标签中学习任务。我们提出在常用的语音重建损失基础上增加两个额外项。第一项构成正则化损失,用于减轻训练语音中扭曲构音的影响;第二项为一致性损失,鼓励生成语音中的正确构音。这些额外损失项由原始和生成语音的帧级构音得分获得,该得分通过单独训练的音素分类器推导得出。在真实舌癌病例上的实验结果表明,合成语音实现了与无障碍自然语音相当的构音质量,同时有效保持了目标说话人的个性。音频样本参见https://myspeechproject.github.io/ArticulationRepair/。