Classification networks can be used to localize and segment objects in images by means of class activation maps (CAMs). However, without pixel-level annotations, classification networks are known to (1) mainly focus on discriminative regions, and (2) to produce diffuse CAMs without well-defined prediction contours. In this work, we approach both problems with two contributions for improving CAM learning. First, we incorporate importance sampling based on the class-wise probability mass function induced by the CAMs to produce stochastic image-level class predictions. This results in CAMs which activate over a larger extent of objects. Second, we formulate a feature similarity loss term which aims to match the prediction contours with edges in the image. As a third contribution, we conduct experiments on the PASCAL VOC 2012 benchmark dataset to demonstrate that these modifications significantly increase the performance in terms of contour accuracy, while being comparable to current state-of-the-art methods in terms of region similarity.
翻译:分类网络可通过类激活图(CAM)对图像中的目标进行定位与分割。然而,在缺乏像素级标注的情况下,分类网络已知存在以下问题:(1)主要关注判别性区域;(2)产生无明确预测轮廓的弥散性CAM。本文针对这两个问题,提出两项改进CAM学习的贡献。首先,我们引入基于CAM诱导的类概率质量函数的重要性采样,以生成随机性图像级类别预测,从而使CAM能够激活目标更广的区域。其次,我们构建一个特征相似性损失项,旨在使预测轮廓与图像边缘相匹配。作为第三项贡献,我们在PASCAL VOC 2012基准数据集上开展实验,证明这些改进在轮廓精度方面显著提升性能,同时与当前最先进方法在区域相似度上具有可比性。