A creative idea is often born from transforming, combining, and modifying ideas from existing visual examples capturing various concepts. However, one cannot simply copy the concept as a whole, and inspiration is achieved by examining certain aspects of the concept. Hence, it is often necessary to separate a concept into different aspects to provide new perspectives. In this paper, we propose a method to decompose a visual concept, represented as a set of images, into different visual aspects encoded in a hierarchical tree structure. We utilize large vision-language models and their rich latent space for concept decomposition and generation. Each node in the tree represents a sub-concept using a learned vector embedding injected into the latent space of a pretrained text-to-image model. We use a set of regularizations to guide the optimization of the embedding vectors encoded in the nodes to follow the hierarchical structure of the tree. Our method allows to explore and discover new concepts derived from the original one. The tree provides the possibility of endless visual sampling at each node, allowing the user to explore the hidden sub-concepts of the object of interest. The learned aspects in each node can be combined within and across trees to create new visual ideas, and can be used in natural language sentences to apply such aspects to new designs.
翻译:创意构思往往源于对蕴含多种概念的现有视觉实例进行转化、组合与修改。然而,直接照搬整体概念并不可行,真正的灵感来源于对概念特定维度的审视。因此,将概念拆解为不同层面以提供新视角显得尤为重要。本文提出一种方法,将视觉概念(以图像集合形式呈现)分解为编码于层级树结构中的不同视觉维度。我们借助大型视觉-语言模型及其丰富的潜在空间实现概念分解与生成。树中每个节点通过注入预训练文本到图像模型的潜在空间的学习向量嵌入表示子概念,并通过一组正则化约束引导节点编码嵌入向量的优化,使其遵循树形层级结构。该方法能够探索并发现源于原始概念的新概念。树结构允许对每个节点进行无限视觉采样,使用户得以探索目标对象隐藏的子概念。各节点学习到的维度可在树内或跨树组合以生成新视觉创意,并可嵌入自然语言语句中,将此类维度应用于新设计。