Polyps are early cancer indicators, so assessing occurrences of polyps and their removal is critical. They are observed through a colonoscopy screening procedure that generates a stream of video frames. Segmenting polyps in their natural video screening procedure has several challenges, such as the co-existence of imaging artefacts, motion blur, and floating debris. Most existing polyp segmentation algorithms are developed on curated still image datasets that do not represent real-world colonoscopy. Their performance often degrades on video data. We propose a video polyp segmentation method that performs self-supervised learning as an auxiliary task and a spatial-temporal self-attention mechanism for improved representation learning. Our end-to-end configuration and joint optimisation of losses enable the network to learn more discriminative contextual features in videos. Our experimental results demonstrate an improvement with respect to several state-of-the-art (SOTA) methods. Our ablation study also confirms that the choice of the proposed joint end-to-end training improves network accuracy by over 3% and nearly 10% on both the Dice similarity coefficient and intersection-over-union compared to the recently proposed method PNS+ and Polyp-PVT, respectively. Results on previously unseen video data indicate that the proposed method generalises.
翻译:息肉是癌症的早期征兆,因此评估息肉的发生并予以切除至关重要。息肉通过结肠镜检查程序进行观察,该程序会生成连续的视频帧序列。在自然视频检查过程中分割息肉面临诸多挑战,例如成像伪影、运动模糊和漂浮碎屑共存。现有的大多数息肉分割算法都是在经过筛选的静态图像数据集上开发的,这些数据集无法代表真实的结肠镜检查环境,导致其在视频数据上的性能往往下降。我们提出了一种视频息肉分割方法,该方法将自监督学习作为辅助任务,并采用时空自注意力机制以改进表征学习。我们的端到端配置与损失函数的联合优化使网络能够学习视频中更具判别力的上下文特征。实验结果表明,相较于多种先进方法,我们的方法取得了性能提升。消融研究也证实,与近期提出的PNS+和Polyp-PVT方法相比,所提出的联合端到端训练方案在Dice相似系数和交并比指标上分别将网络精度提高了超过3%和近10%。在未见过的视频数据上的结果表明,所提方法具有良好的泛化能力。