Language-vision models like CLIP have made significant progress in zero-shot vision tasks, such as zero-shot image classification (ZSIC). However, generating specific and expressive class descriptions remains a major challenge. Existing approaches suffer from granularity and label ambiguity issues. To tackle these challenges, we propose V-GLOSS: Visual Glosses, a novel method leveraging modern language models and semantic knowledge bases to produce visually-grounded class descriptions. We demonstrate V-GLOSS's effectiveness by achieving state-of-the-art results on benchmark ZSIC datasets including ImageNet and STL-10. In addition, we introduce a silver dataset with class descriptions generated by V-GLOSS, and show its usefulness for vision tasks. We make available our code and dataset.
翻译:像CLIP这样的语言-视觉模型在零样本视觉任务(如零样本图像分类)中取得了显著进展。然而,生成具体且富有表达力的类别描述仍是一项重大挑战。现有方法存在粒度不足和标签歧义的问题。为解决这些挑战,我们提出V-GLOSS:视觉注释,一种利用现代语言模型和语义知识库生成视觉感知类别描述的新方法。通过在包括ImageNet和STL-10在内的基准零样本图像分类数据集上取得最新成果,我们证明了V-GLOSS的有效性。此外,我们引入了一个包含由V-GLOSS生成的类别描述的银标准数据集,并展示了其对视觉任务的实用性。我们公开了代码和数据集。