The explosion of short videos has dramatically reshaped the manners people socialize, yielding a new trend for daily sharing and access to the latest information. These rich video resources, on the one hand, benefited from the popularization of portable devices with cameras, but on the other, they can not be independent of the valuable editing work contributed by numerous video creators. In this paper, we investigate a novel and practical problem, namely audio beat matching (ABM), which aims to recommend the proper transition time stamps based on the background music. This technique helps to ease the labor-intensive work during video editing, saving energy for creators so that they can focus more on the creativity of video content. We formally define the ABM problem and its evaluation protocol. Meanwhile, a large-scale audio dataset, i.e., the AutoMatch with over 87k finely annotated background music, is presented to facilitate this newly opened research direction. To further lay solid foundations for the following study, we also propose a novel model termed BeatX to tackle this challenging task. Alongside, we creatively present the concept of label scope, which eliminates the data imbalance issues and assigns adaptive weights for the ground truth during the training procedure in one stop. Though plentiful short video platforms have flourished for a long time, the relevant research concerning this scenario is not sufficient, and to the best of our knowledge, AutoMatch is the first large-scale dataset to tackle the audio beat matching problem. We hope the released dataset and our competitive baseline can encourage more attention to this line of research. The dataset and codes will be made publicly available.
翻译:短视频的爆炸式发展极大地重塑了人们的社交方式,催生了日常分享和获取最新信息的新趋势。这些丰富的视频资源一方面得益于搭载摄像头的便携设备的普及,另一方面也离不开众多视频创作者贡献的宝贵编辑工作。在本文中,我们研究了一个新颖且实用的问题,即音频节拍匹配(ABM),其目标是根据背景音乐推荐合适的过渡时间戳。该技术有助于减轻视频编辑中的劳动密集型工作,为创作者节省精力,使其能更专注于视频内容的创意性。我们正式定义了ABM问题及其评估协议。同时,我们提出一个大规模音频数据集AutoMatch,包含超过8.7万首精细标注的背景音乐,以推动这一新研究方向的发展。为进一步为后续研究奠定坚实基础,我们还提出了一种名为BeatX的新型模型来应对这一挑战性任务。此外,我们创新性地提出了标签范围(label scope)概念,该概念消除了数据不平衡问题,并在训练过程中一次性为真实标签分配自适应权重。尽管短视频平台早已蓬勃发展,但针对这一场景的相关研究仍不充分,据我们所知,AutoMatch是首个解决音频节拍匹配问题的大规模数据集。我们希望发布的数据集及具有竞争力的基线模型能够吸引更多研究者关注这一研究方向。数据集和代码将公开提供。