In the field of personalized image generation, the ability to create images preserving concepts has significantly improved. Creating an image that naturally integrates multiple concepts in a cohesive and visually appealing composition can indeed be challenging. This paper introduces "InstantFamily," an approach that employs a novel masked cross-attention mechanism and a multimodal embedding stack to achieve zero-shot multi-ID image generation. Our method effectively preserves ID as it utilizes global and local features from a pre-trained face recognition model integrated with text conditions. Additionally, our masked cross-attention mechanism enables the precise control of multi-ID and composition in the generated images. We demonstrate the effectiveness of InstantFamily through experiments showing its dominance in generating images with multi-ID, while resolving well-known multi-ID generation problems. Additionally, our model achieves state-of-the-art performance in both single-ID and multi-ID preservation. Furthermore, our model exhibits remarkable scalability with a greater number of ID preservation than it was originally trained with.
翻译:在个性化图像生成领域,生成保留概念图像的能力已显著提升。然而,创建一幅能自然融合多个概念并形成连贯且视觉上令人满意的构图仍具有挑战性。本文提出"InstantFamily"方法,通过新颖的掩码交叉注意力机制与多模态嵌入堆栈,实现零样本多身份图像生成。该方法利用预训练人脸识别模型中的全局与局部特征,结合文本条件,有效保留身份信息。此外,所提出的掩码交叉注意力机制能够精确控制生成图像中的多身份与构图。实验证明,InstantFamily在生成多身份图像方面具有优势,同时解决了已知的多身份生成问题。此外,我们的模型在单身份与多身份保留任务中均达到最先进性能,并且展现出显著的扩展性,能够保留超出原始训练数量的更多身份。