Fine-grained open-set recognition (FineOSR) aims to recognize images belonging to classes with subtle appearance differences while rejecting images of unknown classes. A recent trend in OSR shows the benefit of generative models to discriminative unknown detection. As a type of generative model, energy-based models (EBM) are the potential for hybrid modeling of generative and discriminative tasks. However, most existing EBMs suffer from density estimation in high-dimensional space, which is critical to recognizing images from fine-grained classes. In this paper, we explore the low-dimensional latent space with energy-based prior distribution for OSR in a fine-grained visual world. Specifically, based on the latent space EBM, we propose an attribute-aware information bottleneck (AIB), a residual attribute feature aggregation (RAFA) module, and an uncertainty-based virtual outlier synthesis (UVOS) module to improve the expressivity, granularity, and density of the samples in fine-grained classes, respectively. Our method is flexible to take advantage of recent vision transformers for powerful visual classification and generation. The method is validated on both fine-grained and general visual classification datasets while preserving the capability of generating photo-realistic fake images with high resolution.
翻译:细粒度开放集识别旨在识别具有细微外观差异的类别图像,同时拒绝对未知类别的图像。开放集识别领域的最新趋势显示,生成模型对判别式未知检测具有辅助作用。作为一种生成模型,能量模型具有混合建模生成式与判别式任务的潜力。然而,现有大多数能量模型在高维空间中的密度估计存在困难,这对于识别细粒度类别图像至关重要。本文探索了具有能量先验分布的低维潜在空间,以实现细粒度视觉世界中的开放集识别。具体而言,基于潜在空间能量模型,我们提出了属性感知信息瓶颈、残差属性特征聚合模块及基于不确定性的虚拟离群点合成模块,分别提升细粒度类别的表达能力、粒度和样本密度。该方法可灵活利用近期视觉Transformer实现强大的视觉分类与生成。实验表明,该方法在细粒度与通用视觉分类数据集上均有效,并具备生成高分辨率逼真图像的能力。