Progress in lighting estimation is tracked by computing existing image quality assessment (IQA) metrics on images from standard datasets. While this may appear to be a reasonable approach, we demonstrate that doing so does not correlate to human preference when the estimated lighting is used to relight a virtual scene into a real photograph. To study this, we design a controlled psychophysical experiment where human observers must choose their preference amongst rendered scenes lit using a set of lighting estimation algorithms selected from the recent literature, and use it to analyse how these algorithms perform according to human perception. Then, we demonstrate that none of the most popular IQA metrics from the literature, taken individually, correctly represent human perception. Finally, we show that by learning a combination of existing IQA metrics, we can more accurately represent human preference. This provides a new perceptual framework to help evaluate future lighting estimation algorithms.
翻译:照明估计的进展通常通过计算标准数据集图像上的现有图像质量评估(IQA)指标来衡量。尽管这看似合理,但我们证明,当使用估计的照明将虚拟场景重新光照到真实照片中时,这种方法与人类偏好并不相关。为此,我们设计了一项受控的心理物理实验,要求人类观察者在由近期文献中选出的多种照明估计算法生成的光照渲染场景中做出偏好选择,并以此分析这些算法相对于人类感知的表现。随后,我们发现,文献中最常用的IQA指标单独使用时,均无法准确反映人类感知。最后,我们证明,通过学习组合现有IQA指标,能够更准确地表征人类偏好。这为未来评估照明估计算法提供了新的感知框架。