Recent advances in foundation models present new opportunities for interpretable visual recognition -- one can first query Large Language Models (LLMs) to obtain a set of attributes that describe each class, then apply vision-language models to classify images via these attributes. Pioneering work shows that querying thousands of attributes can achieve performance competitive with image features. However, our further investigation on 8 datasets reveals that LLM-generated attributes in a large quantity perform almost the same as random words. This surprising finding suggests that significant noise may be present in these attributes. We hypothesize that there exist subsets of attributes that can maintain the classification performance with much smaller sizes, and propose a novel learning-to-search method to discover those concise sets of attributes. As a result, on the CUB dataset, our method achieves performance close to that of massive LLM-generated attributes (e.g., 10k attributes for CUB), yet using only 32 attributes in total to distinguish 200 bird species. Furthermore, our new paradigm demonstrates several additional benefits: higher interpretability and interactivity for humans, and the ability to summarize knowledge for a recognition task.
翻译:近期基础模型的研究进展为可解释的视觉识别带来了新机遇——可首先查询大型语言模型以获取描述每个类别的一组属性,随后应用视觉语言模型通过这些属性对图像进行分类。开创性研究表明,查询数千个属性可达到与图像特征相媲美的性能。然而,我们在8个数据集上的进一步研究发现,大规模生成的属性与随机词汇的效果几乎相同。这一惊人发现表明这些属性中可能存在显著噪声。我们假设存在能够以更小规模维持分类性能的属性子集,并提出了一种新颖的学习搜索方法来发现这些简洁的属性集合。因此,在CUB数据集上,我们的方法在仅使用32个属性区分200种鸟类的情况下,性能接近大规模生成的属性(例如CUB的10k个属性)。此外,我们的新范式还展现出若干额外优势:对人类而言更高的可解释性和交互性,以及为识别任务总结知识的能力。