Many deep learning models have achieved dominant performance on the offline beat tracking task. However, online beat tracking, in which only the past and present input features are available, still remains challenging. In this paper, we propose BEAt tracking Streaming Transformer (BEAST), an online joint beat and downbeat tracking system based on the streaming Transformer. To deal with online scenarios, BEAST applies contextual block processing in the Transformer encoder. Moreover, we adopt relative positional encoding in the attention layer of the streaming Transformer encoder to capture relative timing position which is critically important information in music. Carrying out beat and downbeat experiments on benchmark datasets for a low latency scenario with maximum latency under 50 ms, BEAST achieves an F1-measure of 80.04% in beat and 46.78% in downbeat, which is a substantial improvement of about 5 percentage points over the state-of-the-art online beat tracking model.
翻译:许多深度学习模型在离线节拍跟踪任务中已取得显著性能。然而,仅依赖过去和当前输入特征的在线节拍跟踪仍具有挑战性。本文提出基于流式Transformer的节拍跟踪系统BEAST(节拍跟踪流式Transformer),实现节拍与强拍的在线联合跟踪。为适应在线场景,BEAST在Transformer编码器中采用上下文分块处理机制。此外,我们在流式Transformer编码器的注意力层引入相对位置编码,以捕捉音乐中至关重要的相对时序位置信息。在低延迟场景(最大延迟低于50毫秒)下,基于基准数据集的节拍与强拍实验表明:BEAST的节拍F1值达80.04%,强拍F1值达46.78%,较现有最优在线节拍跟踪模型提升约5个百分点。