Despite the recent increase in research on artificial intelligence for music, prominent correlations between key components of lyrics and rhythm such as keywords, stressed syllables, and strong beats are not frequently studied. This is likely due to challenges such as audio misalignment, inaccuracies in syllabic identification, and most importantly, the need for cross-disciplinary knowledge. To address this lack of research, we propose a novel multimodal lyrics-rhythm matching approach in this paper that specifically matches key components of lyrics and music with each other without any language limitations. We use audio instead of sheet music with readily available metadata, which creates more challenges yet increases the application flexibility of our method. Furthermore, our approach creatively generates several patterns involving various multimodalities, including music strong beats, lyrical syllables, auditory changes in a singer's pronunciation, and especially lyrical keywords, which are utilized for matching key lyrical elements with key rhythmic elements. This advantageous approach not only provides a unique way to study auditory lyrics-rhythm correlations including efficient rhythm-based audio alignment algorithms, but also bridges computational linguistics with music as well as music cognition. Our experimental results reveal an 0.81 probability of matching on average, and around 30% of the songs have a probability of 0.9 or higher of keywords landing on strong beats, including 12% of the songs with a perfect landing. Also, the similarity metrics are used to evaluate the correlation between lyrics and rhythm. It shows that nearly 50% of the songs have 0.70 similarity or higher. In conclusion, our approach contributes significantly to the lyrics-rhythm relationship by computationally unveiling insightful correlations.
翻译:尽管近期人工智能音乐研究不断增长,但歌词与节奏关键要素(如关键词、重读音节和强拍)之间的显著关联却鲜少被研究。这主要源于音频对齐偏差、音节识别不准确等挑战,更关键的是需要跨学科知识储备。为填补这一研究空白,本文提出一种新颖的多模态歌词-节奏匹配方法,该方法能无语言限制地实现歌词与音乐关键要素的精准对应。我们选用音频而非附有现成元数据的乐谱作为数据源,虽增加技术难度,却提升了方法的适用灵活性。此外,本方法创造性生成包含多种模态的若干模式——涵盖音乐强拍、歌词音节、歌手发音的听觉变化,特别是歌词关键词——用于实现歌词关键要素与节奏关键要素的匹配。这一优势方法不仅为研究听觉层面的歌词-节奏关联(包括基于节奏的高效音频对齐算法)提供了独特路径,更架起了计算语言学与音乐学乃至音乐认知学的桥梁。实验结果显示,平均匹配概率达0.81,约30%的歌曲中关键词落在强拍的概率超过0.9,其中12%的歌曲实现完美匹配。同时,通过相似度指标评估歌词与节奏的相关性发现,近50%的歌曲相似度达到0.70以上。结论表明,本方法通过计算揭示了富有洞察力的歌词-节奏关联,为该领域研究做出重要贡献。