This paper introduces FALL-E, a foley synthesis system and its training/inference strategies. The FALL-E model employs a cascaded approach comprising low-resolution spectrogram generation, spectrogram super-resolution, and a vocoder. We trained every sound-related model from scratch using our extensive datasets, and utilized a pre-trained language model. We conditioned the model with dataset-specific texts, enabling it to learn sound quality and recording environment based on text input. Moreover, we leveraged external language models to improve text descriptions of our datasets and performed prompt engineering for quality, coherence, and diversity. FALL-E was evaluated by an objective measure as well as listening tests in the DCASE 2023 challenge Task 7. The submission achieved the second place on average, while achieving the best score for diversity, second place for audio quality, and third place for class fitness.
翻译:本文介绍FALL-E,一种拟声音效合成系统及其训练/推理策略。FALL-E模型采用级联方法,包括低分辨率语谱图生成、语谱图超分辨率重建以及声码器。我们使用自建大规模数据集从头训练所有与声音相关的模型,并利用预训练语言模型。通过特定数据集文本对模型进行条件控制,使其能够根据文本输入学习音质和录音环境。此外,我们利用外部语言模型改进数据集文本描述,并针对质量、连贯性和多样性进行提示工程优化。在DCASE 2023挑战赛任务7中,通过客观指标和听力测试对FALL-E进行评估。该提交方案在平均分上位列第二,同时在多样性指标上获得最佳分数,音频质量位列第二,类别匹配度位列第三。