This survey presents a comprehensive analysis of data augmentation techniques in human-centric vision tasks, a first of its kind in the field. It delves into a wide range of research areas including person ReID, human parsing, human pose estimation, and pedestrian detection, addressing the significant challenges posed by overfitting and limited training data in these domains. Our work categorizes data augmentation methods into two main types: data generation and data perturbation. Data generation covers techniques like graphic engine-based generation, generative model-based generation, and data recombination, while data perturbation is divided into image-level and human-level perturbations. Each method is tailored to the unique requirements of human-centric tasks, with some applicable across multiple areas. Our contributions include an extensive literature review, providing deep insights into the influence of these augmentation techniques in human-centric vision and highlighting the nuances of each method. We also discuss open issues and future directions, such as the integration of advanced generative models like Latent Diffusion Models, for creating more realistic and diverse training data. This survey not only encapsulates the current state of data augmentation in human-centric vision but also charts a course for future research, aiming to develop more robust, accurate, and efficient human-centric vision systems.
翻译:本综述对以人为中心的视觉任务中的数据增强技术进行了全面分析,是该领域内首次此类研究。它深入探讨了包括行人重识别、人体解析、人体姿态估计和行人检测在内的广泛研究领域,针对这些领域中过拟合和训练数据有限带来的重大挑战。我们的工作将数据增强方法分为两大类:数据生成和数据扰动。数据生成涵盖基于图形引擎的生成、基于生成模型的生成及数据重组等技术,而数据扰动则分为图像级和人体级扰动。每种方法均针对以人为中心的任务的独特需求进行定制,部分方法可跨多个领域适用。我们的贡献包括广泛的文献综述,深入揭示了这些增强技术在以人为中心视觉中的影响,并突出了每种方法的细微差别。我们还探讨了开放性问题与未来方向,例如整合先进生成模型(如潜在扩散模型)以创建更真实、更多样的训练数据。本综述不仅总结了以人为中心视觉中数据增强的现状,还为未来研究指明了方向,旨在开发更稳健、准确、高效的以人为中心的视觉系统。