3D human body shape and pose estimation from RGB images is a challenging problem with potential applications in augmented/virtual reality, healthcare and fitness technology and virtual retail. Recent solutions have focused on three types of inputs: i) single images, ii) multi-view images and iii) videos. In this study, we surveyed and compared 3D body shape and pose estimation methods for contemporary dance and performing arts, with a special focus on human body pose and dressing, camera viewpoint, illumination conditions and background conditions. We demonstrated that multi-frame methods, such as PHALP, provide better results than single-frame method for pose estimation when dancers are performing contemporary dances.
翻译:从RGB图像中估计三维人体体态与姿态是一个具有挑战性的问题,在增强/虚拟现实、医疗与健身技术以及虚拟零售领域具有潜在应用前景。近期解决方案主要集中于三类输入:i) 单张图像,ii) 多视角图像,以及iii) 视频。本研究针对当代舞蹈与表演艺术,系统调研并比较了三维人体体态与姿态估计方法,特别关注人体姿态与着装、相机视角、光照条件及背景条件等因素的影响。研究表明,当舞者表演当代舞蹈时,PHALP等多帧方法在姿态估计方面优于单帧方法。