Interpretable representation learning has been playing a key role in creative intelligent systems. In the music domain, current learning algorithms can successfully learn various features such as pitch, timbre, chord, texture, etc. However, most methods rely heavily on music domain knowledge. It remains an open question what general computational principles give rise to interpretable representations, especially low-dim factors that agree with human perception. In this study, we take inspiration from modern physics and use physical symmetry as a self-consistency constraint for the latent space. Specifically, it requires the prior model that characterises the dynamics of the latent states to be equivariant with respect to certain group transformations. We show that physical symmetry leads the model to learn a linear pitch factor from unlabelled monophonic music audio in a self-supervised fashion. In addition, the same methodology can be applied to computer vision, learning a 3D Cartesian space from videos of a simple moving object without labels. Furthermore, physical symmetry naturally leads to representation augmentation, a new technique which improves sample efficiency.
翻译:可解释表示学习在创造性智能系统中一直扮演着关键角色。在音乐领域,当前学习算法能够成功学习音高、音色、和弦、织体等多种特征。然而,大多数方法严重依赖音乐领域知识。何种通用计算原理能催生符合人类感知的可解释表示(尤其是低维因子)仍是一个开放性问题。本研究从现代物理学汲取灵感,将物理对称性作为潜在空间的自我一致性约束。具体而言,该方法要求刻画潜在状态动态特性的先验模型在特定群变换下具有等变性。我们证明,物理对称性促使模型以自监督方式从未标记的单声道音乐音频中学习线性音高因子。此外,同一方法论可应用于计算机视觉领域,通过无标签的简单运动物体视频学习三维笛卡尔空间。更具意义的是,物理对称性自然引出了表征增强这一技术,该方法能有效提升样本效率。