Many applications of cross-modal music retrieval are related to connecting sheet music images to audio recordings. A typical and recent approach to this is to learn, via deep neural networks, a joint embedding space that correlates short fixed-size snippets of audio and sheet music by means of an appropriate similarity structure. However, two challenges that arise out of this strategy are the requirement of strongly aligned data to train the networks, and the inherent discrepancies of musical content between audio and sheet music snippets caused by local and global tempo differences. In this paper, we address these two shortcomings by designing a cross-modal recurrent network that learns joint embeddings that can summarize longer passages of corresponding audio and sheet music. The benefits of our method are that it only requires weakly aligned audio-sheet music pairs, as well as that the recurrent network handles the non-linearities caused by tempo variations between audio and sheet music. We conduct a number of experiments on synthetic and real piano data and scores, showing that our proposed recurrent method leads to more accurate retrieval in all possible configurations.
翻译:跨模态音乐检索的许多应用涉及将乐谱图像与音频录音相关联。近期一种典型方法是利用深度神经网络学习一个联合嵌入空间,通过适当的相似度结构关联音频与乐谱的短固定长度片段。然而,该策略面临两个挑战:训练网络需要强对齐数据,以及局部与全局速度差异导致音频与乐谱片段之间的音乐内容固有偏差。本文通过设计跨模态循环网络解决这两个缺陷,该网络学习能够总结较长对应音频与乐谱段落的联合嵌入。本方法的优势在于仅需弱对齐的音频-乐谱对,同时循环网络能够处理由音频与乐谱间速度变化引起的非线性问题。我们在合成与真实钢琴数据及乐谱上开展多项实验,结果表明所提出的循环方法在所有配置下均能实现更精确的检索。