Colonoscopic Polyp Re-Identification aims to match a specific polyp in a large gallery with different cameras and views, which plays a key role for the prevention and treatment of colorectal cancer in the computer-aided diagnosis. However, traditional methods mainly focus on the visual representation learning, while neglect to explore the potential of semantic features during training, which may easily leads to poor generalization capability when adapted the pretrained model into the new scenarios. To relieve this dilemma, we propose a simple but effective training method named VT-ReID, which can remarkably enrich the representation of polyp videos with the interchange of high-level semantic information. Moreover, we elaborately design a novel clustering mechanism to introduce prior knowledge from textual data, which leverages contrastive learning to promote better separation from abundant unlabeled text data. To the best of our knowledge, this is the first attempt to employ the visual-text feature with clustering mechanism for the colonoscopic polyp re-identification. Empirical results show that our method significantly outperforms current state-of-the art methods with a clear margin.
翻译:结肠镜息肉重识别旨在通过不同相机和视角在大规模图库中匹配特定息肉,这对计算机辅助诊断中结直肠癌的预防与治疗至关重要。然而,传统方法主要关注视觉表征学习,忽视了训练过程中语义特征的潜在价值,这容易导致预训练模型在新场景下泛化能力不足。为解决这一困境,我们提出了一种简单而有效的训练方法VT-ReID,通过高级语义信息的交互显著丰富了息肉视频的表征。此外,我们精心设计了一种新型聚类机制,从文本数据中引入先验知识,利用对比学习促进从丰富的无标签文本数据中实现更好的分离。据我们所知,这是首次将带聚类机制的视觉-文本特征应用于结肠镜息肉重识别。实验结果表明,我们的方法以显著优势超越了当前最优方法。