Manifold learning flows are a class of generative modelling techniques that assume a low-dimensional manifold description of the data. The embedding of such a manifold into the high-dimensional space of the data is achieved via learnable invertible transformations. Therefore, once the manifold is properly aligned via a reconstruction loss, the probability density is tractable on the manifold and maximum likelihood can be used to optimize the network parameters. Naturally, the lower-dimensional representation of the data requires an injective-mapping. Recent approaches were able to enforce that the density aligns with the modelled manifold, while efficiently calculating the density volume-change term when embedding to the higher-dimensional space. However, unless the injective-mapping is analytically predefined, the learned manifold is not necessarily an efficient representation of the data. Namely, the latent dimensions of such models frequently learn an entangled intrinsic basis, with degenerate information being stored in each dimension. Alternatively, if a locally orthogonal and/or sparse basis is to be learned, here coined canonical intrinsic basis, it can serve in learning a more compact latent space representation. Toward this end, we propose a canonical manifold learning flow method, where a novel optimization objective enforces the transformation matrix to have few prominent and non-degenerate basis functions. We demonstrate that by minimizing the off-diagonal manifold metric elements $\ell_1$-norm, we can achieve such a basis, which is simultaneously sparse and/or orthogonal. Canonical manifold flow yields a more efficient use of the latent space, automatically generating fewer prominent and distinct dimensions to represent data, and a better approximation of target distributions than other manifold flow methods in most experiments we conducted, resulting in lower FID scores.
翻译:流形学习流是一类假设数据具有低维流形描述的生成建模技术。通过可学习的可逆变换,将此类流形嵌入到数据的高维空间中。因此,一旦通过重构损失正确对齐流形,流形上的概率密度便可处理,并可用最大似然法优化网络参数。自然地,数据的低维表示需要单射映射。近期方法能够在嵌入到高维空间时有效计算密度体积变化项,同时确保密度与建模流形对齐。然而,除非单射映射以解析方式预定义,否则学习到的流形未必是数据的高效表示。即,此类模型的潜在维度常学习到纠缠的固有基,导致每个维度存储退化信息。若学习局部正交和/或稀疏基(此处称为规范固有基),则有助于学习更紧凑的潜在空间表示。为此,我们提出一种规范流形学习流方法,其新型优化目标强制变换矩阵具有少量显著且非退化的基函数。我们证明,通过最小化流形度量矩阵的非对角元素$\ell_1$范数,可实现同时具有稀疏性和/或正交性的此类基。与其它流形流方法相比,规范流形流在大多实验中更高效地利用潜在空间,自动生成更少的显著区分维度表示数据,更优近似目标分布,从而获得更低的FID分数。