While abundant research has been conducted on improving high-level visual understanding and reasoning capabilities of large multimodal models~(LMMs), their visual quality assessment~(IQA) ability has been relatively under-explored. Here we take initial steps towards this goal by employing the two-alternative forced choice~(2AFC) prompting, as 2AFC is widely regarded as the most reliable way of collecting human opinions of visual quality. Subsequently, the global quality score of each image estimated by a particular LMM can be efficiently aggregated using the maximum a posterior estimation. Meanwhile, we introduce three evaluation criteria: consistency, accuracy, and correlation, to provide comprehensive quantifications and deeper insights into the IQA capability of five LMMs. Extensive experiments show that existing LMMs exhibit remarkable IQA ability on coarse-grained quality comparison, but there is room for improvement on fine-grained quality discrimination. The proposed dataset sheds light on the future development of IQA models based on LMMs. The codes will be made publicly available at https://github.com/h4nwei/2AFC-LMMs.
翻译:尽管已有大量研究致力于提升大型多模态模型在高层视觉理解与推理能力方面的表现,但这些模型在视觉质量评估方面的能力却相对未得到充分探索。为此,我们率先采用二选一强制选择提示方法(2AFC)来推进该目标,因为2AFC被广泛认为是收集人类对视觉质量意见的最可靠方式。随后,通过最大后验概率估计,可高效聚合特定LMM对每幅图像所评估的全局质量分数。同时,我们引入三个评估准则:一致性、准确性和相关性,以对五种LMM的图像质量评估能力进行量化综合评估与深度分析。大量实验表明,现有LMM在粗粒度质量比较任务中展现出卓越的IQA能力,但在细粒度质量判别方面仍有提升空间。本文提出的数据集将为基于LMM的IQA模型未来发展提供启示。相关代码将在https://github.com/h4nwei/2AFC-LMMs 公开。