The analysis of vision-based deep neural networks (DNNs) is highly desirable but it is very challenging due to the difficulty of expressing formal specifications for vision tasks and the lack of efficient verification procedures. In this paper, we propose to leverage emerging multimodal, vision-language, foundation models (VLMs) as a lens through which we can reason about vision models. VLMs have been trained on a large body of images accompanied by their textual description, and are thus implicitly aware of high-level, human-understandable concepts describing the images. We describe a logical specification language $\texttt{Con}_{\texttt{spec}}$ designed to facilitate writing specifications in terms of these concepts. To define and formally check $\texttt{Con}_{\texttt{spec}}$ specifications, we build a map between the internal representations of a given vision model and a VLM, leading to an efficient verification procedure of natural-language properties for vision models. We demonstrate our techniques on a ResNet-based classifier trained on the RIVAL-10 dataset using CLIP as the multimodal model.
翻译:对基于视觉的深度神经网络(DNNs)进行分析非常必要,但由于难以表达视觉任务的形式化规范且缺乏高效的验证流程,这一工作极具挑战性。本文提出利用新兴的多模态、视觉-语言基础模型(VLMs)作为分析视觉模型的透镜。VLMs经过大量图像及其文本描述的训练,因此隐式地理解描述图像的高层、人类可理解概念。我们设计了一种逻辑规范语言$\texttt{Con}_{\texttt{spec}}$,旨在便于基于这些概念编写规范。为定义并形式化检验$\texttt{Con}_{\texttt{spec}}$规范,我们在给定视觉模型的内部表示与VLM之间建立映射,从而形成一种针对视觉模型自然语言属性的高效验证流程。我们基于ResNet分类器(在RIVAL-10数据集上训练)并采用CLIP作为多模态模型,对所提技术进行了验证。