StyleGAN's disentangled style representation enables powerful image editing by manipulating the latent variables, but accurately mapping real-world images to their latent variables (GAN inversion) remains a challenge. Existing GAN inversion methods struggle to maintain editing directions and produce realistic results. To address these limitations, we propose Make It So, a novel GAN inversion method that operates in the $\mathcal{Z}$ (noise) space rather than the typical $\mathcal{W}$ (latent style) space. Make It So preserves editing capabilities, even for out-of-domain images. This is a crucial property that was overlooked in prior methods. Our quantitative evaluations demonstrate that Make It So outperforms the state-of-the-art method PTI~\cite{roich2021pivotal} by a factor of five in inversion accuracy and achieves ten times better edit quality for complex indoor scenes.
翻译:StyleGAN的解耦风格表示通过操控潜变量实现了强大的图像编辑能力,但将真实图像精确映射到其潜变量(GAN反演)仍是一项挑战。现有GAN反演方法难以保持编辑方向并产生逼真结果。为克服这些局限,我们提出Make It So——一种在$\mathcal{Z}$(噪声)空间而非典型$\mathcal{W}$(潜在风格)空间运行的新型GAN反演方法。Make It So能够保持编辑能力,即使对于域外图像亦是如此。这一关键特性在先前方法中被忽视。定量评估表明,Make It So在反演精度上以五倍优势超越现有最优方法PTI~\cite{roich2021pivotal},且在复杂室内场景中实现十倍于后者的编辑质量。