We have recently seen great progress in learning interpretable music representations, ranging from basic factors, such as pitch and timbre, to high-level concepts, such as chord and texture. However, most methods rely heavily on music domain knowledge. It remains an open question what general computational principles give rise to interpretable representations, especially low-dim factors that agree with human perception. In this study, we take inspiration from modern physics and use physical symmetry as a self consistency constraint for the latent space of time-series data. Specifically, it requires the prior model that characterises the dynamics of the latent states to be equivariant with respect to certain group transformations. We show that physical symmetry leads the model to learn a linear pitch factor from unlabelled monophonic music audio in a self-supervised fashion. In addition, the same methodology can be applied to computer vision, learning a 3D Cartesian space from videos of a simple moving object without labels. Furthermore, physical symmetry naturally leads to counterfactual representation augmentation, a new technique which improves sample efficiency.
翻译:我们近期在学习可解释的音乐表示方面取得了重要进展,从音高、音色等基本要素到和弦、织体等高级概念。然而,大多数方法严重依赖音乐领域知识。如何通过通用的计算原理获得可解释的表示(特别是与人类感知一致的低维因素)仍是一个悬而未决的问题。在本研究中,我们从现代物理学中汲取灵感,将物理对称性作为时序数据潜空间的自我一致性约束。具体而言,该方法要求描述潜状态动态特性的先验模型对特定群变换满足等变性。研究表明,物理对称性能够使模型以自监督方式从未标记的单声道音乐音频中学习线性音高因子。此外,该方法同样适用于计算机视觉领域——无需标签即可从简单运动物体的视频中学习三维笛卡尔空间。进一步地,物理对称性自然衍生出反事实表示增强这一新技术,该技术能有效提升样本效率。