Medical image segmentation is fundamental to clinical decision-making, yet existing models remain fragmented. They are usually trained on single knowledge sources and specific to individual tasks, modalities, or organs. This fragmentation contrasts sharply with clinical practice, where experts seamlessly integrate diverse knowledge: anatomical priors from training, exemplar-based reasoning from reference cases, and iterative refinement through real-time interaction. We present $\textbf{K-Prism}$, a unified segmentation framework that mirrors this clinical flexibility by systematically integrating three knowledge paradigms: (i) $\textit{semantic priors}$ learned from annotated datasets, (ii) $\textit{in-context knowledge}$ from few-shot reference examples, and (iii) $\textit{interactive feedback}$ from user inputs like clicks or scribbles. Our key insight is that these heterogeneous knowledge sources can be encoded into a dual-prompt representation: 1-D sparse prompts defining $\textit{what}$ to segment and 2-D dense prompts indicating $\textit{where}$ to attend, which are then dynamically routed through a Mixture-of-Experts (MoE) decoder. This design enables flexible switching between paradigms and joint training across diverse tasks without architectural modifications. Comprehensive experiments on 18 public datasets spanning diverse modalities (CT, MRI, X-ray, pathology, ultrasound, etc.) demonstrate that K-Prism achieves state-of-the-art performance across semantic, in-context, and interactive segmentation settings.
翻译:医学图像分割是临床决策的基础,然而现有模型仍存在碎片化问题。它们通常基于单一知识源训练,且仅适用于特定任务、模态或器官。这种碎片化与临床实践形成鲜明对比:在临床中,专家能无缝整合多样化知识——训练中习得的解剖先验、基于参考案例的范例推理以及通过实时交互实现的迭代优化。我们提出$\textbf{K-Prism}$统一分割框架,通过系统整合三种知识范式来模拟这种临床灵活性:(i)从标注数据集中习得的$\textit{语义先验}$,(ii)来自少样本参考样例的$\textit{上下文知识}$,以及(iii)来自点击或涂鸦等用户输入的$\textit{交互反馈}$。我们的关键洞察在于,这些异质知识源可编码为双提示表征:定义$\textit{分割目标}$的1维稀疏提示和指示$\textit{关注区域}$的2维密集提示,两者随后通过混合专家(MoE)解码器动态路由。该设计支持不同范式间的灵活切换,无需修改架构即可实现跨任务的联合训练。在涵盖CT、MRI、X光、病理学、超声等多种模态的18个公开数据集上的综合实验表明,K-Prism在语义、上下文和交互式分割场景中均达到最先进性能。