Human-object interactions (HOIs) are crucial for human-centric scene understanding applications such as human-centric visual generation, AR/VR, and robotics. Since existing methods mainly explore capturing HOIs, rendering HOI remains less investigated. In this paper, we address this challenge in HOI animation from a compositional perspective, i.e., animating novel HOIs including novel interaction, novel human and/or novel object driven by a novel pose sequence. Specifically, we adopt neural human-object deformation to model and render HOI dynamics based on implicit neural representations. To enable the interaction pose transferring among different persons and objects, we then devise a new compositional conditional neural radiance field (or CC-NeRF), which decomposes the interdependence between human and object using latent codes to enable compositionally animation control of novel HOIs. Experiments show that the proposed method can generalize well to various novel HOI animation settings. Our project page is https://zhihou7.github.io/CHONA/
翻译:人-物交互(HOI)对于以人为中心的场景理解应用(如以人为中心的视觉生成、增强现实/虚拟现实和机器人技术)至关重要。由于现有方法主要探索捕捉HOI,HOI的渲染仍然研究不足。本文从组合视角应对HOI动画中的这一挑战,即对新颖的HOI进行动画生成,包括由新颖姿态序列驱动的新颖交互、新颖人体和/或新颖物体。具体而言,我们采用神经人体-物体变形技术,基于隐式神经表示建模并渲染HOI动态。为实现不同人体与物体之间的交互姿态迁移,我们进一步设计了一种新颖的组合式条件神经辐射场(简称CC-NeRF),通过潜码分解人与物体之间的相互依赖关系,从而实现对新颖HOI的组合式动画控制。实验表明,所提方法能够良好泛化至多种新颖HOI动画场景。项目页面为https://zhihou7.github.io/CHONA/