The rapid increase in user-generated-content (UGC) videos calls for the development of effective video quality assessment (VQA) algorithms. However, the objective of the UGC-VQA problem is still ambiguous and can be viewed from two perspectives: the technical perspective, measuring the perception of distortions; and the aesthetic perspective, which relates to preference and recommendation on contents. To understand how these two perspectives affect overall subjective opinions in UGC-VQA, we conduct a large-scale subjective study to collect human quality opinions on overall quality of videos as well as perceptions from aesthetic and technical perspectives. The collected Disentangled Video Quality Database (DIVIDE-3k) confirms that human quality opinions on UGC videos are universally and inevitably affected by both aesthetic and technical perspectives. In light of this, we propose the Disentangled Objective Video Quality Evaluator (DOVER) to learn the quality of UGC videos based on the two perspectives. The DOVER proves state-of-the-art performance in UGC-VQA under very high efficiency. With perspective opinions in DIVIDE-3k, we further propose DOVER++, the first approach to provide reliable clear-cut quality evaluations from a single aesthetic or technical perspective. Code at https://github.com/VQAssessment/DOVER.
翻译:用户生成内容(UGC)视频的快速增长推动了有效视频质量评估(VQA)算法的发展。然而,UGC-VQA问题的目标仍不明确,可从两个视角审视:技术视角衡量对失真的感知;美学视角则涉及内容偏好与推荐。为理解这两个视角如何影响UGC-VQA中的整体主观意见,我们开展大规模主观研究,收集人类对视频整体质量以及美学与技术视角的感知评价。所构建的解耦视频质量数据库(DIVIDE-3k)证实,人类对UGC视频的质量意见普遍且不可避免地受到美学与技术视角的双重影响。基于此,我们提出解耦客观视频质量评估器(DOVER),从两个视角学习UGC视频质量。DOVER在极高效率下展现了UGC-VQA领域的最先进性能。借助DIVIDE-3k中的视角意见,我们进一步提出DOVER++,这是首个能可靠提供单一美学或技术视角清晰质量评估的方法。代码见https://github.com/VQAssessment/DOVER。