Learning disentangled representations in an unsupervised manner is a fundamental challenge in machine learning. Solving it may unlock other problems, such as generalization, interpretability, or fairness. While remarkably difficult to solve in general, recent works have shown that disentanglement is provably achievable under additional assumptions that can leverage geometrical constraints, such as local isometry. To use these insights, we propose a novel perspective on disentangled representation learning built on quadratic optimal transport. Specifically, we formulate the problem in the Gromov-Monge setting, which seeks isometric mappings between distributions supported on different spaces. We propose the Gromov-Monge-Gap (GMG), a regularizer that quantifies the geometry-preservation of an arbitrary push-forward map between two distributions supported on different spaces. We demonstrate the effectiveness of GMG regularization for disentanglement on four standard benchmarks. Moreover, we show that geometry preservation can even encourage unsupervised disentanglement without the standard reconstruction objective - making the underlying model decoder-free, and promising a more practically viable and scalable perspective on unsupervised disentanglement.
翻译:无监督学习中的解耦表示学习是机器学习中的一个基础性挑战。解决该问题可能为泛化性、可解释性或公平性等其他问题提供突破口。尽管该问题在一般情况下极难求解,但近期研究表明,在能够利用几何约束(如局部等距性)的附加假设下,解耦在理论上是可以实现的。基于这些见解,我们提出了一种建立在二次最优传输理论基础上的解耦表示学习新视角。具体而言,我们将问题置于Gromov-Monge框架中,该框架寻求定义在不同空间上的分布之间的等距映射。我们提出了Gromov-Monge间隙(GMG)——一种正则化器,用于量化定义在不同空间上的两个分布之间任意前推映射的几何保持性。我们在四个标准基准测试上验证了GMG正则化对于解耦的有效性。此外,我们发现几何保持性甚至可以在不使用标准重构目标的情况下促进无监督解耦——这使得底层模型无需解码器,为无监督解耦提供了一个更具实践可行性和可扩展性的前景。