Deep learning based methods have become a paradigm for cover song identification (CSI) in recent years, where the ByteCover systems have achieved state-of-the-art results on all the mainstream datasets of CSI. However, with the burgeon of short videos, many real-world applications require matching short music excerpts to full-length music tracks in the database, which is still under-explored and waiting for an industrial-level solution. In this paper, we upgrade the previous ByteCover systems to ByteCover3 that utilizes local features to further improve the identification performance of short music queries. ByteCover3 is designed with a local alignment loss (LAL) module and a two-stage feature retrieval pipeline, allowing the system to perform CSI in a more precise and efficient way. We evaluated ByteCover3 on multiple datasets with different benchmark settings, where ByteCover3 beat all the compared methods including its previous versions.
翻译:基于深度学习的方法近年来已成为翻唱歌曲识别的主流范式,其中ByteCover系统在所有主流CSI数据集上均取得了最先进的结果。然而,随着短视频的蓬勃发展,许多实际应用场景需要将短音乐片段与数据库中的完整音乐曲目进行匹配,这一领域尚未得到充分探索,亟待工业级解决方案。本文对先前的ByteCover系统进行升级,提出ByteCover3系统,通过利用局部特征进一步提升短音乐查询的识别性能。ByteCover3设计了局部对齐损失模块与两阶段特征检索流程,使系统能够更精确高效地执行CSI任务。我们在不同基准设置下的多个数据集上对ByteCover3进行评估,结果表明ByteCover3超越了包括其先前版本在内的所有对比方法。