Recent successes in image synthesis are powered by large-scale diffusion models. However, most methods are currently limited to either text- or image-conditioned generation for synthesizing an entire image, texture transfer or inserting objects into a user-specified region. In contrast, in this work we focus on synthesizing complex interactions (ie, an articulated hand) with a given object. Given an RGB image of an object, we aim to hallucinate plausible images of a human hand interacting with it. We propose a two-step generative approach: a LayoutNet that samples an articulation-agnostic hand-object-interaction layout, and a ContentNet that synthesizes images of a hand grasping the object given the predicted layout. Both are built on top of a large-scale pretrained diffusion model to make use of its latent representation. Compared to baselines, the proposed method is shown to generalize better to novel objects and perform surprisingly well on out-of-distribution in-the-wild scenes of portable-sized objects. The resulting system allows us to predict descriptive affordance information, such as hand articulation and approaching orientation. Project page: https://judyye.github.io/affordiffusion-www
翻译:近期图像合成的成功得益于大规模扩散模型。然而,当前大多数方法仅局限于基于文本或图像条件的全图合成、纹理迁移或对象插入至用户指定区域。与此不同,本研究聚焦于合成给定对象上的复杂交互(即铰接式手部)。给定对象的RGB图像,我们旨在生成人类手部与之交互的合理图像。我们提出了一种两步生成方法:LayoutNet用于采样与关节无关的手-物交互布局,ContentNet则基于预测布局合成手部抓取对象的图像。两者均构建于大规模预训练扩散模型之上,以利用其潜在表示。与基线方法相比,所提方法对新颖对象表现出更强的泛化能力,并在便携尺寸对象的分布外真实场景中取得显著效果。该系统能够预测描述性功能信息,如手部关节构型和接近方向。项目页面:https://judyye.github.io/affordiffusion-www