We consider the problem of generating musical soundtracks in sync with rhythmic visual cues. Most existing works rely on pre-defined music representations, leading to the incompetence of generative flexibility and complexity. Other methods directly generating video-conditioned waveforms suffer from limited scenarios, short lengths, and unstable generation quality. To this end, we present Long-Term Rhythmic Video Soundtracker (LORIS), a novel framework to synthesize long-term conditional waveforms. Specifically, our framework consists of a latent conditional diffusion probabilistic model to perform waveform synthesis. Furthermore, a series of context-aware conditioning encoders are proposed to take temporal information into consideration for a long-term generation. Notably, we extend our model's applicability from dances to multiple sports scenarios such as floor exercise and figure skating. To perform comprehensive evaluations, we establish a benchmark for rhythmic video soundtracks including the pre-processed dataset, improved evaluation metrics, and robust generative baselines. Extensive experiments show that our model generates long-term soundtracks with state-of-the-art musical quality and rhythmic correspondence. Codes are available at \url{https://github.com/OpenGVLab/LORIS}.
翻译:我们研究了与节奏视觉线索同步生成音乐配乐的问题。现有工作大多依赖预定义的音乐表征,导致生成灵活性和复杂性不足。其他直接生成视频条件波形的方法则受限于场景单一、时长较短和生成质量不稳定。为此,我们提出了长期节奏视频配乐器(LORIS)——一个用于合成长期条件波形的新型框架。具体而言,我们的框架包含一个潜在条件扩散概率模型,用于执行波形合成。此外,一系列上下文感知条件编码器被提出以考虑时间信息,实现长期生成。值得注意的是,我们将模型的应用场景从舞蹈扩展至自由体操和花样滑冰等多项体育场景。为进行全面评估,我们建立了一个节奏视频配乐基准,包括预处理数据集、改进的评估指标和稳健的生成基线。大量实验表明,我们的模型在音乐质量和节奏对应性方面生成了具有最先进水平的长期配乐。代码可在 \url{https://github.com/OpenGVLab/LORIS} 获取。