Perceptual image quality assessment (IQA) is the task of predicting the visual quality of an image as perceived by a human observer. Current state-of-the-art techniques are based on deep representations trained in discriminative manner. Such representations may ignore visually important features, if they are not predictive of class labels. Recent generative models successfully learn low-dimensional representations using auto-encoding and have been argued to preserve better visual features. Here we leverage existing auto-encoders and propose VAE-QA, a simple and efficient method for predicting image quality in the presence of a full-reference. We evaluate our approach on four standard benchmarks and find that it significantly improves generalization across datasets, has fewer trainable parameters, a smaller memory footprint and faster run time.
翻译:感知图像质量评估(IQA)是预测图像被人类观察者感知到的视觉质量的任务。当前最先进的技术基于以判别方式训练的深度表示。这种表示可能忽略视觉上重要的特征——如果这些特征对类别标签没有预测能力的话。最近的生成模型通过自编码成功学习了低维表示,并被论证能更好地保留视觉特征。本文利用现有自编码器提出了VAE-QA,一种在全参考条件下预测图像质量的简单高效方法。我们在四个标准基准上评估了该方法,发现其在跨数据集泛化方面显著提升,同时具有更少的可训练参数、更小的内存占用和更快的运行时间。