We consider the problem of generating musical soundtracks in sync with rhythmic visual cues. Most existing works rely on pre-defined music representations, leading to the incompetence of generative flexibility and complexity. Other methods directly generating video-conditioned waveforms suffer from limited scenarios, short lengths, and unstable generation quality. To this end, we present Long-Term Rhythmic Video Soundtracker (LORIS), a novel framework to synthesize long-term conditional waveforms. Specifically, our framework consists of a latent conditional diffusion probabilistic model to perform waveform synthesis. Furthermore, a series of context-aware conditioning encoders are proposed to take temporal information into consideration for a long-term generation. Notably, we extend our model's applicability from dances to multiple sports scenarios such as floor exercise and figure skating. To perform comprehensive evaluations, we establish a benchmark for rhythmic video soundtracks including the pre-processed dataset, improved evaluation metrics, and robust generative baselines. Extensive experiments show that our model generates long-term soundtracks with state-of-the-art musical quality and rhythmic correspondence. Codes are available at \url{https://github.com/OpenGVLab/LORIS}.


翻译:我们考虑了生成与节奏性视觉线索同步的音乐配乐问题。现有大多数工作依赖预定义的音乐表示,导致生成灵活性和复杂性不足。其他直接生成视频条件波形的方法则受限于场景有限、时长较短以及生成质量不稳定。为此,我们提出了长时节奏视频配乐器(LORIS),这是一种用于合成长时间条件波形的新型框架。具体而言,我们的框架包含一个潜在条件扩散概率模型用于波形合成。此外,我们设计了一系列上下文感知条件编码器,将时间信息纳入考虑以实现长时间生成。值得注意的是,我们将模型的应用场景从舞蹈扩展到了多种体育场景,如自由体操和花样滑冰。为了进行全面评估,我们建立了一个节奏视频配乐基准,包括预处理数据集、改进的评估指标以及鲁棒的生成基线。大量实验表明,我们的模型生成的长时间配乐在音乐质量和节奏对应性方面达到了最先进的水平。代码可在 \url{https://github.com/OpenGVLab/LORIS} 获取。

0
下载
关闭预览

相关内容

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集
专知会员服务
167+阅读 · 2020年3月18日
强化学习最新教程,17页pdf
专知会员服务
182+阅读 · 2019年10月11日
[综述]深度学习下的场景文本检测与识别
专知会员服务
78+阅读 · 2019年10月10日
Hierarchically Structured Meta-learning
CreateAMind
27+阅读 · 2019年5月22日
Transferring Knowledge across Learning Processes
CreateAMind
29+阅读 · 2019年5月18日
强化学习的Unsupervised Meta-Learning
CreateAMind
18+阅读 · 2019年1月7日
Unsupervised Learning via Meta-Learning
CreateAMind
44+阅读 · 2019年1月3日
A Technical Overview of AI & ML in 2018 & Trends for 2019
待字闺中
18+阅读 · 2018年12月24日
【跟踪Tracking】15篇论文+代码 | 中秋快乐~
专知
18+阅读 · 2018年9月24日
disentangled-representation-papers
CreateAMind
26+阅读 · 2018年9月12日
【推荐】GAN架构入门综述(资源汇总)
机器学习研究会
10+阅读 · 2017年9月3日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2013年12月31日
国家自然科学基金
0+阅读 · 2013年12月31日
国家自然科学基金
0+阅读 · 2013年12月31日
国家自然科学基金
1+阅读 · 2012年12月31日
国家自然科学基金
0+阅读 · 2012年12月31日
国家自然科学基金
0+阅读 · 2011年12月31日
国家自然科学基金
0+阅读 · 2009年12月31日
国家自然科学基金
0+阅读 · 2009年12月31日
Arxiv
0+阅读 · 2023年7月19日
Arxiv
17+阅读 · 2021年3月29日
VIP会员
最新内容
综述 | 终端智能体:命令行环境中的 AI Agents
专知会员服务
2+阅读 · 8月24日
综述 | 多模态智能体框架基础与前沿
专知会员服务
1+阅读 · 8月24日
《自适应无人机集群网络开发》
专知会员服务
4+阅读 · 8月24日
无人机与作战飞机最佳精确目标定位技术
专知会员服务
3+阅读 · 8月24日
博士论文 | 大动作空间中的在线与离线策略学习
论文 | 全球负责任 AI 指数 2026 方法论
专知会员服务
7+阅读 · 8月23日
超致命战场空间中的战术通信生存能力
专知会员服务
5+阅读 · 8月23日
综述 | 面向机器人的视觉触觉智能
专知会员服务
6+阅读 · 8月22日
论文 | 基础模型驱动具身智能体安全综述
专知会员服务
4+阅读 · 8月22日
相关基金
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2013年12月31日
国家自然科学基金
0+阅读 · 2013年12月31日
国家自然科学基金
0+阅读 · 2013年12月31日
国家自然科学基金
1+阅读 · 2012年12月31日
国家自然科学基金
0+阅读 · 2012年12月31日
国家自然科学基金
0+阅读 · 2011年12月31日
国家自然科学基金
0+阅读 · 2009年12月31日
国家自然科学基金
0+阅读 · 2009年12月31日
Top
微信扫码咨询专知VIP会员