In musical compositions that include vocals, lyrics significantly contribute to artistic expression. Consequently, previous studies have introduced the concept of a recommendation system that suggests lyrics similar to a user's favorites or personalized preferences, aiding in the discovery of lyrics among millions of tracks. However, many of these systems do not fully consider human perceptions of lyric similarity, primarily due to limited research in this area. To bridge this gap, we conducted a comparative analysis of computational methods for modeling lyric similarity with human perception. Results indicated that computational models based on similarities between embeddings from pre-trained BERT-based models, the audio from which the lyrics are derived, and phonetic components are indicative of perceptual lyric similarity. This finding underscores the importance of semantic, stylistic, and phonetic similarities in human perception about lyric similarity. We anticipate that our findings will enhance the development of similarity-based lyric recommendation systems by offering pseudo-labels for neural network development and introducing objective evaluation metrics.
翻译:在包含人声的音乐作品中,歌词对艺术表达贡献显著。因此,先前研究引入了推荐系统的概念,该系统能够推荐与用户喜好或个性化偏好相似的歌词,帮助用户在数百万曲目中发掘歌词。然而,由于该领域研究有限,许多系统并未充分考虑人类对歌词相似性的感知。为弥补这一空白,我们开展了歌词相似性建模计算方法与人类感知的对比分析。结果表明,基于预训练BERT模型嵌入相似性、歌词所源自的音频特征以及语音成分的计算模型,能够指示感知层面的歌词相似性。这一发现强调了语义、风格及语音相似性在人类歌词感知中的重要性。我们预期,本研究将为基于相似性的歌词推荐系统提供神经网络开发的伪标签及客观评价指标,从而推动其发展。