Automatic high-quality rendering of anime scenes from complex real-world images is of significant practical value. The challenges of this task lie in the complexity of the scenes, the unique features of anime style, and the lack of high-quality datasets to bridge the domain gap. Despite promising attempts, previous efforts are still incompetent in achieving satisfactory results with consistent semantic preservation, evident stylization, and fine details. In this study, we propose Scenimefy, a novel semi-supervised image-to-image translation framework that addresses these challenges. Our approach guides the learning with structure-consistent pseudo paired data, simplifying the pure unsupervised setting. The pseudo data are derived uniquely from a semantic-constrained StyleGAN leveraging rich model priors like CLIP. We further apply segmentation-guided data selection to obtain high-quality pseudo supervision. A patch-wise contrastive style loss is introduced to improve stylization and fine details. Besides, we contribute a high-resolution anime scene dataset to facilitate future research. Our extensive experiments demonstrate the superiority of our method over state-of-the-art baselines in terms of both perceptual quality and quantitative performance.
翻译:摘要:从复杂真实世界图像自动生成高质量的动漫场景具有重要的实际价值。该任务的挑战在于场景的复杂性、动漫风格的独特特征以及缺乏高质量数据集来弥合领域差距。尽管已有一些有前景的尝试,但先前的研究在实现一致语义保留、显著风格化和精细细节的令人满意结果方面仍显不足。在本研究中,我们提出了Scenimefy,一种新颖的半监督图像到图像翻译框架以应对这些挑战。我们的方法利用结构一致的伪配对数据指导学习,简化了纯无监督设置。这些伪数据独特地源自一个受语义约束的StyleGAN,该模型利用了丰富的模型先验知识(如CLIP)。我们还应用了基于分割的数据选择来获取高质量的伪监督。我们引入了一种基于块对比的风格损失,以改进风格化和精细细节。此外,我们贡献了一个高分辨率的动漫场景数据集,以促进未来研究。我们的大量实验证明了我们的方法在感知质量和量化性能方面均优于最先进的基线方法。