Large-scale text-to-image models including Stable Diffusion are capable of generating high-fidelity photorealistic portrait images. There is an active research area dedicated to personalizing these models, aiming to synthesize specific subjects or styles using provided sets of reference images. However, despite the plausible results from these personalization methods, they tend to produce images that often fall short of realism and are not yet on a commercially viable level. This is particularly noticeable in portrait image generation, where any unnatural artifact in human faces is easily discernible due to our inherent human bias. To address this, we introduce MagiCapture, a personalization method for integrating subject and style concepts to generate high-resolution portrait images using just a few subject and style references. For instance, given a handful of random selfies, our fine-tuned model can generate high-quality portrait images in specific styles, such as passport or profile photos. The main challenge with this task is the absence of ground truth for the composed concepts, leading to a reduction in the quality of the final output and an identity shift of the source subject. To address these issues, we present a novel Attention Refocusing loss coupled with auxiliary priors, both of which facilitate robust learning within this weakly supervised learning setting. Our pipeline also includes additional post-processing steps to ensure the creation of highly realistic outputs. MagiCapture outperforms other baselines in both quantitative and qualitative evaluations and can also be generalized to other non-human objects.
翻译:大型文本到图像模型(包括Stable Diffusion)能够生成高保真的逼真肖像图像。当前研究领域积极致力于个性化这些模型,旨在利用提供的参考图像集合合成特定主体或风格。然而,尽管这些个性化方法产生了看似合理的结果,但其生成的图像往往缺乏真实感,尚未达到商业可用水平。这一点在肖像图像生成中尤为明显,由于人类固有的认知偏差,人脸上的任何不自然痕迹都容易被察觉。为解决此问题,我们提出MagiCapture,一种通过少量主体和风格参考图像整合主体与风格概念、生成高分辨率肖像图像的个性化方法。例如,仅凭几张随机自拍,我们微调后的模型即可生成特定风格(如证件照或头像照)的高质量肖像。该任务的主要挑战在于合成概念缺乏真实标注,导致最终输出质量下降及源主体身份偏移。为此,我们提出新颖的注意力重聚焦损失(Attention Refocusing loss)结合辅助先验,两者在弱监督学习框架下促进稳健学习。我们的流程还包含额外后处理步骤,以确保生成高度逼真的输出。在定量与定性评估中,MagiCapture均优于其他基线方法,并可泛化至非人类物体。