Despite the recent increase in research on artificial intelligence for music, prominent correlations between key components of lyrics and rhythm such as keywords, stressed syllables, and strong beats are not frequently studied. This is likely due to challenges such as audio misalignment, inaccuracies in syllabic identification, and most importantly, the need for cross-disciplinary knowledge. To address this lack of research, we propose a novel multimodal lyrics-rhythm matching approach in this paper that specifically matches key components of lyrics and music with each other without any language limitations. We use audio instead of sheet music with readily available metadata, which creates more challenges yet increases the application flexibility of our method. Furthermore, our approach creatively generates several patterns involving various multimodalities, including music strong beats, lyrical syllables, auditory changes in a singer's pronunciation, and especially lyrical keywords, which are utilized for matching key lyrical elements with key rhythmic elements. This advantageous approach not only provides a unique way to study auditory lyrics-rhythm correlations including efficient rhythm-based audio alignment algorithms, but also bridges computational linguistics with music as well as music cognition. Our experimental results reveal an 0.81 probability of matching on average, and around 30% of the songs have a probability of 0.9 or higher of keywords landing on strong beats, including 12% of the songs with a perfect landing. Also, the similarity metrics are used to evaluate the correlation between lyrics and rhythm. It shows that nearly 50% of the songs have 0.70 similarity or higher. In conclusion, our approach contributes significantly to the lyrics-rhythm relationship by computationally unveiling insightful correlations.


翻译:尽管近期人工智能音乐研究不断增长,但歌词与节奏关键要素(如关键词、重读音节和强拍)之间的显著关联却鲜少被研究。这主要源于音频对齐偏差、音节识别不准确等挑战,更关键的是需要跨学科知识储备。为填补这一研究空白,本文提出一种新颖的多模态歌词-节奏匹配方法,该方法能无语言限制地实现歌词与音乐关键要素的精准对应。我们选用音频而非附有现成元数据的乐谱作为数据源,虽增加技术难度,却提升了方法的适用灵活性。此外,本方法创造性生成包含多种模态的若干模式——涵盖音乐强拍、歌词音节、歌手发音的听觉变化,特别是歌词关键词——用于实现歌词关键要素与节奏关键要素的匹配。这一优势方法不仅为研究听觉层面的歌词-节奏关联(包括基于节奏的高效音频对齐算法)提供了独特路径,更架起了计算语言学与音乐学乃至音乐认知学的桥梁。实验结果显示,平均匹配概率达0.81,约30%的歌曲中关键词落在强拍的概率超过0.9,其中12%的歌曲实现完美匹配。同时,通过相似度指标评估歌词与节奏的相关性发现,近50%的歌曲相似度达到0.70以上。结论表明,本方法通过计算揭示了富有洞察力的歌词-节奏关联,为该领域研究做出重要贡献。

0
下载
关闭预览

相关内容

专知会员服务
124+阅读 · 2020年9月8日
100+篇《自监督学习(Self-Supervised Learning)》论文最新合集
专知会员服务
167+阅读 · 2020年3月18日
[综述]深度学习下的场景文本检测与识别
专知会员服务
78+阅读 · 2019年10月10日
机器学习入门的经验与建议
专知会员服务
94+阅读 · 2019年10月10日
【SIGGRAPH2019】TensorFlow 2.0深度学习计算机图形学应用
专知会员服务
41+阅读 · 2019年10月9日
Hierarchically Structured Meta-learning
CreateAMind
27+阅读 · 2019年5月22日
IEEE | DSC 2019诚邀稿件 (EI检索)
Call4Papers
10+阅读 · 2019年2月25日
逆强化学习-学习人先验的动机
CreateAMind
16+阅读 · 2019年1月18日
强化学习的Unsupervised Meta-Learning
CreateAMind
18+阅读 · 2019年1月7日
Unsupervised Learning via Meta-Learning
CreateAMind
44+阅读 · 2019年1月3日
A Technical Overview of AI & ML in 2018 & Trends for 2019
待字闺中
18+阅读 · 2018年12月24日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2013年12月31日
国家自然科学基金
0+阅读 · 2012年12月31日
国家自然科学基金
0+阅读 · 2012年12月31日
Arxiv
0+阅读 · 2023年5月5日
Arxiv
29+阅读 · 2023年1月12日
Arxiv
31+阅读 · 2021年6月30日
Arxiv
10+阅读 · 2017年12月29日
Arxiv
151+阅读 · 2017年8月1日
VIP会员
最新内容
俄乌无人机战争的六大启示
专知会员服务
9+阅读 · 8月3日
《无人机空中监控:通信实验洞察》
专知会员服务
6+阅读 · 8月3日
从采集到决策:美军视角下的战术情报范式重构
《履带式无人地面战车技术发展现状》
专知会员服务
6+阅读 · 8月2日
《无人机脆弱性利用:网络空间力量的新域》
专知会员服务
9+阅读 · 8月1日
美空军如何将人工智能从战场部署至后方机关
专知会员服务
14+阅读 · 7月31日
Top
微信扫码咨询专知VIP会员