Research in Image Generation has recently made significant progress, particularly boosted by the introduction of Vision-Language models which are able to produce high-quality visual content based on textual inputs. Despite ongoing advancements in terms of generation quality and realism, no methodical frameworks have been defined yet to quantitatively measure the quality of the generated content and the adherence with the prompted requests: so far, only human-based evaluations have been adopted for quality satisfaction and for comparing different generative methods. We introduce a novel automated method for Visual Concept Evaluation (ViCE), i.e. to assess consistency between a generated/edited image and the corresponding prompt/instructions, with a process inspired by the human cognitive behaviour. ViCE combines the strengths of Large Language Models (LLMs) and Visual Question Answering (VQA) into a unified pipeline, aiming to replicate the human cognitive process in quality assessment. This method outlines visual concepts, formulates image-specific verification questions, utilizes the Q&A system to investigate the image, and scores the combined outcome. Although this brave new hypothesis of mimicking humans in the image evaluation process is in its preliminary assessment stage, results are promising and open the door to a new form of automatic evaluation which could have significant impact as the image generation or the image target editing tasks become more and more sophisticated.
翻译:图像生成研究近期取得了显著进展,尤其是视觉语言模型的引入极大推动了基于文本输入生成高质量视觉内容的发展。尽管在生成质量和逼真度方面持续进步,但目前尚未建立系统化框架来定量衡量生成内容的质量及其与提示请求的一致性:迄今为止,质量满意度评估及不同生成方法的比较仍仅依赖人工评价。我们提出了一种创新的自动化视觉概念评估(ViCE)方法,即通过模拟人类认知行为的过程,评估生成/编辑图像与对应提示/指令之间的一致性。ViCE将大语言模型(LLMs)与视觉问答(VQA)的优势整合至统一流程中,旨在复现人类在质量评估中的认知过程。该方法通过提取视觉概念、生成图像特异性验证问题、利用问答系统分析图像,并对综合结果进行评分。尽管这种模仿人类进行图像评估的大胆新假设仍处于初步评估阶段,但已有成果展现了良好前景,为新型自动化评估开辟了道路——随着图像生成或目标图像编辑任务日益复杂化,该评估方法可能产生重大影响。