Segment Anything Model (SAM) is a foundation model for semantic segmentation and shows excellent generalization capability with the prompts. In this empirical study, we investigate the robustness and zero-shot generalizability of the SAM in the domain of robotic surgery in various settings of (i) prompted vs. unprompted; (ii) bounding box vs. points-based prompt; (iii) generalization under corruptions and perturbations with five severity levels; and (iv) state-of-the-art supervised model vs. SAM. We conduct all the observations with two well-known robotic instrument segmentation datasets of MICCAI EndoVis 2017 and 2018 challenges. Our extensive evaluation results reveal that although SAM shows remarkable zero-shot generalization ability with bounding box prompts, it struggles to segment the whole instrument with point-based prompts and unprompted settings. Furthermore, our qualitative figures demonstrate that the model either failed to predict the parts of the instrument mask (e.g., jaws, wrist) or predicted parts of the instrument as different classes in the scenario of overlapping instruments within the same bounding box or with the point-based prompt. In fact, it is unable to identify instruments in some complex surgical scenarios of blood, reflection, blur, and shade. Additionally, SAM is insufficiently robust to maintain high performance when subjected to various forms of data corruption. Therefore, we can argue that SAM is not ready for downstream surgical tasks without further domain-specific fine-tuning.
翻译:分割一切模型(SAM)是一种用于语义分割的基础模型,在提示引导下展现出卓越的泛化能力。本实证研究从以下四个维度探究SAM在机器人手术领域的鲁棒性与零样本泛化能力:(i)有提示与无提示的对比;(ii)边界框提示与基于点的提示;(iii)五级严重程度下面对数据损坏和扰动的泛化性能;(iv)当前最先进的监督模型与SAM的对比。我们利用MICCAI EndoVis 2017和2018挑战赛中的两个知名机器人手术器械分割数据集开展全部观测实验。大量评估结果表明:尽管SAM在边界框提示下展现出显著的零样本泛化能力,但在基于点的提示及无提示场景中,其难以完整分割手术器械。此外,定性分析显示:当同一边界框内或基于点提示场景中存在器械重叠时,模型要么无法预测器械掩膜的局部区域(如夹爪、腕部),要么将器械部分区域预测为不同类别。事实上,在某些复杂手术场景(如血液、反光、模糊、阴影)中,模型完全无法识别器械。同时,SAM在面对多种数据损坏扰动时,鲁棒性不足,难以维持高性能表现。因此,我们认为SAM在缺乏领域特异性微调的情况下,尚不能胜任下游手术任务。