Morphological atlases are an important tool in organismal studies, and modern high-throughput Computed Tomography (CT) facilities can produce hundreds of full-body high-resolution volumetric images of organisms. However, creating an atlas from these volumes requires accurate organ segmentation. In the last decade, machine learning approaches have achieved incredible results in image segmentation tasks, but they require large amounts of annotated data for training. In this paper, we propose a self-training framework for multi-organ segmentation in tomographic images of Medaka fish. We utilize the pseudo-labeled data from a pretrained Teacher model and adopt a Quality Classifier to refine the pseudo-labeled data. Then, we introduce a pixel-wise knowledge distillation method to prevent overfitting to the pseudo-labeled data and improve the segmentation performance. The experimental results demonstrate that our method improves mean Intersection over Union (IoU) by 5.9% on the full dataset and enables keeping the quality while using three times less markup.
翻译:形态学图谱是生物体研究中的重要工具,现代高通量计算机断层扫描(CT)设施可生成数百张生物体全身高分辨率体积图像。然而,从这些体积数据创建图谱需要精确的器官分割。过去十年中,机器学习方法在图像分割任务中取得了令人瞩目的成果,但这类方法需要大量标注数据进行训练。本文提出了一种用于青鳉鱼断层扫描图像多器官分割的自训练框架。我们利用预训练教师模型生成的伪标签数据,并采用质量分类器对其精炼。随后引入像素级知识蒸馏方法,以防止对伪标签数据的过拟合并提升分割性能。实验结果表明,本方法在全数据集上将平均交并比(IoU)提升了5.9%,且在使用三倍减少标注量的情况下仍能保持分割质量。