Exploring a machine learning system to generate meaningful combinatorial object images from multiple textual descriptions, emulating human creativity, is a significant challenge as humans are able to construct amazing combinatorial objects, but machines strive to emulate data distribution. In this paper, we develop a straightforward yet highly effective technique called acceptable swap-sampling to generate a combinatorial object image that exhibits novelty and surprise, utilizing text concepts of different objects. Initially, we propose a swapping mechanism that constructs a novel embedding by exchanging column vectors of two text embeddings for generating a new combinatorial image through a cutting-edge diffusion model. Furthermore, we design an acceptable region by managing suitable CLIP distances between the new image and the original concept generations, increasing the likelihood of accepting the new image with a high-quality combination. This region allows us to efficiently sample a small subset from a new image pool generated by using randomly exchanging column vectors. Lastly, we employ a segmentation method to compare CLIP distances among the segmented components, ultimately selecting the most promising object image from the sampled subset. Our experiments focus on text pairs of objects from ImageNet, and our results demonstrate that our approach outperforms recent methods such as Stable-Diffusion2, DALLE2, ERNIE-ViLG2 and Bing in generating novel and surprising object images, even when the associated concepts appear to be implausible, such as lionfish-abacus. Furthermore, during the sampling process, our approach without training and human preference is also comparable to PickScore and HPSv2 trained using human preference datasets.
翻译:探索一个机器学习系统,使其能够从多个文本描述中生成有意义的组合性物体图像,并模拟人类创造力,是一项重大挑战:人类能够构建令人惊叹的组合性物体,但机器却难以超越数据分布的模仿。本文提出一种直接而高效的技术——可接受交换采样,通过利用不同物体的文本概念,生成兼具新颖性与惊喜感的组合性物体图像。首先,我们设计了一种交换机制:通过交换两个文本嵌入向量的列向量,构建新型嵌入,并借助先进的扩散模型生成全新的组合图像。其次,我们通过控制新图像与原始概念生成图像之间的合适CLIP距离,设计了一个可接受区域,从而提高接受具有高质量组合特性新图像的可能性。该区域使我们能够从通过随机交换列向量生成的新图像池中高效采样出一个小规模子集。最后,我们采用分割方法比较各分割部件之间的CLIP距离,从采样子集中最终选出最具潜力的物体图像。实验聚焦于ImageNet中成对物体文本,结果显示:即使面对看似不合理的概念组合(如狮子鱼-算盘),我们的方法在生成新颖且令人惊喜的物体图像方面,仍优于Stable-Diffusion2、DALLE2、ERNIE-ViLG2及Bing等最新方法。此外,在采样过程中,无需训练或人类偏好的本方法,其性能也与使用人类偏好数据集训练的PickScore和HPSv2相当。