Visual In-Context Learning (ICL) has emerged as a promising research area due to its capability to accomplish various tasks with limited example pairs through analogical reasoning. However, training-based visual ICL has limitations in its ability to generalize to unseen tasks and requires the collection of a diverse task dataset. On the other hand, existing methods in the inference-based visual ICL category solely rely on textual prompts, which fail to capture fine-grained contextual information from given examples and can be time-consuming when converting from images to text prompts. To address these challenges, we propose Analogist, a novel inference-based visual ICL approach that exploits both visual and textual prompting techniques using a text-to-image diffusion model pretrained for image inpainting. For visual prompting, we propose a self-attention cloning (SAC) method to guide the fine-grained structural-level analogy between image examples. For textual prompting, we leverage GPT-4V's visual reasoning capability to efficiently generate text prompts and introduce a cross-attention masking (CAM) operation to enhance the accuracy of semantic-level analogy guided by text prompts. Our method is out-of-the-box and does not require fine-tuning or optimization. It is also generic and flexible, enabling a wide range of visual tasks to be performed in an in-context manner. Extensive experiments demonstrate the superiority of our method over existing approaches, both qualitatively and quantitatively.
翻译:视觉上下文学习(Visual In-Context Learning, ICL)因其通过类比推理实现有限样本对下的多任务能力,已成为极具前景的研究方向。然而,基于训练的视觉ICL面临泛化至未见任务的局限,且需收集多样化的任务数据集。另一方面,现有推理型视觉ICL方法仅依赖文本提示,既难以捕获给定示例中的细粒度上下文信息,从图像转换至文本提示的过程也耗时较长。为解决上述挑战,我们提出Analogist——一种新型推理型视觉ICL方法,该方法利用预训练用于图像修复的文本到图像扩散模型,同时融合视觉与文本提示技术。在视觉提示方面,我们提出自注意力克隆(Self-Attention Cloning, SAC)方法,用于引导图像示例间的细粒度结构级类比;在文本提示方面,我们借助GPT-4V的视觉推理能力高效生成文本提示,并引入跨注意力掩码(Cross-Attention Masking, CAM)操作以增强文本引导的语义级类比精度。本方法即开即用,无需微调或优化。其通用性与灵活性使其能够以上下文学习方式执行多样化的视觉任务。大量实验从定性与定量两方面均验证了本方法相较于现有技术的优越性。