The impressive success of style-based GANs (StyleGANs) in high-fidelity image synthesis has motivated research to understand the semantic properties of their latent spaces. In this paper, we approach this problem through a geometric analysis of latent spaces as a manifold. In particular, we propose a local dimension estimation algorithm for arbitrary intermediate layers in a pre-trained GAN model. The estimated local dimension is interpreted as the number of possible semantic variations from this latent variable. Moreover, this intrinsic dimension estimation enables unsupervised evaluation of disentanglement for a latent space. Our proposed metric, called Distortion, measures an inconsistency of intrinsic tangent space on the learned latent space. Distortion is purely geometric and does not require any additional attribute information. Nevertheless, Distortion shows a high correlation with the global-basis-compatibility and supervised disentanglement score. Our work is the first step towards selecting the most disentangled latent space among various latent spaces in a GAN without attribute labels.
翻译:基于风格的生成对抗网络(StyleGANs)在高保真图像合成领域取得的显著成功,推动了对其中潜在空间语义属性的研究。本文从流形角度对潜在空间进行几何分析,提出了一种针对预训练GAN中任意中间层的局部维度估计算法。估计得到的局部维度被解释为该潜在变量可能产生的语义变化数量。此外,这种内在维度估计能够实现对潜在空间解缠性的无监督评估。我们提出的指标——失真度(Distortion),用于衡量学习到的潜在空间中固有切空间的不一致性。该度量完全基于几何特性,无需任何额外属性信息。尽管如此,失真度仍与全局基兼容性和有监督解缠性得分呈现高度相关性。本研究首次实现了在无需属性标签的情况下,从GAN的多个潜在空间中筛选出解缠性最优的潜在空间。