AI-Generated Images (AGIs) have inherent multimodal nature. Unlike traditional image quality assessment (IQA) on natural scenarios, AGIs quality assessment (AGIQA) takes the correspondence of image and its textual prompt into consideration. This is coupled in the ground truth score, which confuses the unimodal IQA methods. To solve this problem, we introduce IP-IQA (AGIs Quality Assessment via Image and Prompt), a multimodal framework for AGIQA via corresponding image and prompt incorporation. Specifically, we propose a novel incremental pretraining task named Image2Prompt for better understanding of AGIs and their corresponding textual prompts. An effective and efficient image-prompt fusion module, along with a novel special [QA] token, are also applied. Both are plug-and-play and beneficial for the cooperation of image and its corresponding prompt. Experiments demonstrate that our IP-IQA achieves the state-of-the-art on AGIQA-1k and AGIQA-3k datasets. Code will be available at https://github.com/Coobiw/IP-IQA.
翻译:AI生成图像(AGIs)具有固有的多模态特性。与面向自然场景的传统图像质量评估(IQA)不同,AGIs质量评估(AGIQA)需考虑图像与其文本提示之间的对应关系。这种耦合关系体现在地面真实评分中,使得单模态IQA方法难以有效处理。为解决该问题,我们提出IP-IQA(基于图像与提示的AGIs质量评估),一种通过关联图像与提示进行AGIQA的多模态框架。具体而言,我们设计了一项名为Image2Prompt的新型增量预训练任务,以增强对AGIs及其对应文本提示的理解。此外,还引入了高效且有效的图像-提示融合模块及其新型专用[QA]标记。二者均即插即用,有助于促进图像与其提示的协同作用。实验表明,我们的IP-IQA在AGIQA-1k和AGIQA-3k数据集上达到了最优性能。代码将开源在https://github.com/Coobiw/IP-IQA。