Text-to-audio (TTA) generation is a recent popular problem that aims to synthesize general audio given text descriptions. Previous methods utilized latent diffusion models to learn audio embedding in a latent space with text embedding as the condition. However, they ignored the synchronization between audio and visual content in the video, and tended to generate audio mismatching from video frames. In this work, we propose a novel and personalized text-to-sound generation approach with visual alignment based on latent diffusion models, namely DiffAVA, that can simply fine-tune lightweight visual-text alignment modules with frozen modality-specific encoders to update visual-aligned text embeddings as the condition. Specifically, our DiffAVA leverages a multi-head attention transformer to aggregate temporal information from video features, and a dual multi-modal residual network to fuse temporal visual representations with text embeddings. Then, a contrastive learning objective is applied to match visual-aligned text embeddings with audio features. Experimental results on the AudioCaps dataset demonstrate that the proposed DiffAVA can achieve competitive performance on visual-aligned text-to-audio generation.
翻译:摘要:文本到音频生成是近期广受关注的问题,旨在根据文本描述合成通用音频。现有方法利用潜在扩散模型,以文本嵌入为条件在潜在空间中学习音频嵌入,但忽略了视频中音频与视觉内容的同步性,易生成与视频帧不匹配的音频。为此,本文提出一种新颖的个性化文本到声音生成方法DiffAVA,该方法基于潜在扩散模型实现视觉对齐,通过冻结特定模态编码器并仅微调轻量级视觉文本对齐模块,将视觉对齐后的文本嵌入作为条件。具体而言,DiffAVA采用多头注意力Transformer聚合视频特征中的时序信息,并通过双多模态残差网络融合时序视觉表征与文本嵌入。随后,引入对比学习目标,使视觉对齐后的文本嵌入与音频特征相匹配。在AudioCaps数据集上的实验结果表明,所提出的DiffAVA在视觉对齐的文本到音频生成任务中取得了具有竞争力的性能。