State-of-the-art end-to-end Optical Music Recognition (OMR) has, to date, primarily been carried out using monophonic transcription techniques to handle complex score layouts, such as polyphony, often by resorting to simplifications or specific adaptations. Despite their efficacy, these approaches imply challenges related to scalability and limitations. This paper presents the Sheet Music Transformer, the first end-to-end OMR model designed to transcribe complex musical scores without relying solely on monophonic strategies. Our model employs a Transformer-based image-to-sequence framework that predicts score transcriptions in a standard digital music encoding format from input images. Our model has been tested on two polyphonic music datasets and has proven capable of handling these intricate music structures effectively. The experimental outcomes not only indicate the competence of the model, but also show that it is better than the state-of-the-art methods, thus contributing to advancements in end-to-end OMR transcription.
翻译:迄今为止,最先进的端到端光学乐谱识别(OMR)主要采用单音转录技术处理复杂乐谱布局(如复音),通常通过简化或特定适配来实现。尽管这些方法有效,但存在可扩展性不足与局限性问题。本文提出Sheet Music Transformer——首个不依赖单音策略即可转录复杂乐谱的端到端OMR模型。该模型采用基于Transformer的图像到序列框架,能从输入图像直接预测标准数字音乐编码格式的乐谱转录结果。我们在两个复音乐谱数据集上测试了模型,验证其能有效处理此类复杂音乐结构。实验结果显示,该模型不仅具备卓越能力,更优于现有最优方法,从而推动端到端OMR转录技术的发展。