Humans constantly contact objects to move and perform tasks. Thus, detecting human-object contact is important for building human-centered artificial intelligence. However, there exists no robust method to detect contact between the body and the scene from an image, and there exists no dataset to learn such a detector. We fill this gap with HOT ("Human-Object conTact"), a new dataset of human-object contacts for images. To build HOT, we use two data sources: (1) We use the PROX dataset of 3D human meshes moving in 3D scenes, and automatically annotate 2D image areas for contact via 3D mesh proximity and projection. (2) We use the V-COCO, HAKE and Watch-n-Patch datasets, and ask trained annotators to draw polygons for the 2D image areas where contact takes place. We also annotate the involved body part of the human body. We use our HOT dataset to train a new contact detector, which takes a single color image as input, and outputs 2D contact heatmaps as well as the body-part labels that are in contact. This is a new and challenging task that extends current foot-ground or hand-object contact detectors to the full generality of the whole body. The detector uses a part-attention branch to guide contact estimation through the context of the surrounding body parts and scene. We evaluate our detector extensively, and quantitative results show that our model outperforms baselines, and that all components contribute to better performance. Results on images from an online repository show reasonable detections and generalizability.
翻译:人类通过持续接触物体来移动和执行任务。因此,检测人-物体接触对于构建以人为中心的人工智能至关重要。然而,目前尚无鲁棒方法可从图像中检测人体与场景之间的接触,也不存在可用于训练此类检测器的数据集。我们通过HOT(人-物体接触)数据集填补了这一空白——该数据集专为图像中的人-物体接触而构建。为构建HOT,我们采用两种数据源:(1) 利用PROX数据集中三维场景中运动的三维人体网格数据,通过三维网格邻近度和投影自动标注二维图像区域中的接触;(2) 使用V-COCO、HAKE和Watch-n-Patch数据集,由专业标注员绘制接触区域的二维多边形,并标注涉及的人体部位。基于HOT数据集,我们训练了一种新型接触检测器,该检测器以单张彩色图像为输入,输出二维接触热图及接触部位标签。这是一个全新的挑战性任务,将现有脚-地或手-物接触检测器扩展至全身接触的通用场景。检测器采用部位注意力分支,通过人体周围部位与场景的上下文信息引导接触估计。我们对该检测器进行了全面评估,定量结果表明模型性能优于基线方法,且各组件均有助于提升效果。在线图像库的测试结果验证了检测结果的合理性与泛化能力。