This paper introduces the Global-Local Image Perceptual Score (GLIPS), an image metric designed to assess the photorealistic image quality of AI-generated images with a high degree of alignment to human visual perception. Traditional metrics such as FID and KID scores do not align closely with human evaluations. The proposed metric incorporates advanced transformer-based attention mechanisms to assess local similarity and Maximum Mean Discrepancy (MMD) to evaluate global distributional similarity. To evaluate the performance of GLIPS, we conducted a human study on photorealistic image quality. Comprehensive tests across various generative models demonstrate that GLIPS consistently outperforms existing metrics like FID, SSIM, and MS-SSIM in terms of correlation with human scores. Additionally, we introduce the Interpolative Binning Scale (IBS), a refined scaling method that enhances the interpretability of metric scores by aligning them more closely with human evaluative standards. The proposed metric and scaling approach not only provides more reliable assessments of AI-generated images but also suggest pathways for future enhancements in image generation technologies.
翻译:本文提出全局-局部图像感知分数(GLIPS),这是一种旨在评估AI生成图像逼真质量的图像度量指标,其评估结果与人类视觉感知高度一致。传统的FID和KID等度量指标与人类评价的吻合度较低。该指标结合了基于Transformer的先进注意力机制来评估局部相似性,并采用最大均值差异(MMD)来评估全局分布相似性。为评估GLIPS的性能,我们开展了一项关于逼真图像质量的人工研究。针对多种生成模型的综合测试表明,在与人类评分的相关性方面,GLIPS始终优于FID、SSIM和MS-SSIM等现有指标。此外,我们引入了插值分箱尺度(IBS),这是一种改进的尺度化方法,通过使度量分数更贴近人类评价标准来增强其可解释性。所提出的度量指标与尺度化方法不仅能为AI生成图像提供更可靠的评估,还为图像生成技术的未来改进指明了方向。