We consider the problem of generating musical soundtracks in sync with rhythmic visual cues. Most existing works rely on pre-defined music representations, leading to the incompetence of generative flexibility and complexity. Other methods directly generating video-conditioned waveforms suffer from limited scenarios, short lengths, and unstable generation quality. To this end, we present Long-Term Rhythmic Video Soundtracker (LORIS), a novel framework to synthesize long-term conditional waveforms. Specifically, our framework consists of a latent conditional diffusion probabilistic model to perform waveform synthesis. Furthermore, a series of context-aware conditioning encoders are proposed to take temporal information into consideration for a long-term generation. Notably, we extend our model's applicability from dances to multiple sports scenarios such as floor exercise and figure skating. To perform comprehensive evaluations, we establish a benchmark for rhythmic video soundtracks including the pre-processed dataset, improved evaluation metrics, and robust generative baselines. Extensive experiments show that our model generates long-term soundtracks with state-of-the-art musical quality and rhythmic correspondence. Codes are available at \url{https://github.com/OpenGVLab/LORIS}.
翻译:我们考虑了生成与节奏性视觉线索同步的音乐配乐问题。现有大多数工作依赖预定义的音乐表示,导致生成灵活性和复杂性不足。其他直接生成视频条件波形的方法则受限于场景有限、时长较短以及生成质量不稳定。为此,我们提出了长时节奏视频配乐器(LORIS),这是一种用于合成长时间条件波形的新型框架。具体而言,我们的框架包含一个潜在条件扩散概率模型用于波形合成。此外,我们设计了一系列上下文感知条件编码器,将时间信息纳入考虑以实现长时间生成。值得注意的是,我们将模型的应用场景从舞蹈扩展到了多种体育场景,如自由体操和花样滑冰。为了进行全面评估,我们建立了一个节奏视频配乐基准,包括预处理数据集、改进的评估指标以及鲁棒的生成基线。大量实验表明,我们的模型生成的长时间配乐在音乐质量和节奏对应性方面达到了最先进的水平。代码可在 \url{https://github.com/OpenGVLab/LORIS} 获取。