Direct image-to-graph transformation is a challenging task that solves object detection and relationship prediction in a single model. Due to the complexity of this task, large training datasets are rare in many domains, which makes the training of large networks challenging. This data sparsity necessitates the establishment of pre-training strategies akin to the state-of-the-art in computer vision. In this work, we introduce a set of methods enabling cross-domain and cross-dimension transfer learning for image-to-graph transformers. We propose (1) a regularized edge sampling loss for sampling the optimal number of object relationships (edges) across domains, (2) a domain adaptation framework for image-to-graph transformers that aligns features from different domains, and (3) a simple projection function that allows us to pretrain 3D transformers on 2D input data. We demonstrate our method's utility in cross-domain and cross-dimension experiments, where we pretrain our models on 2D satellite images before applying them to vastly different target domains in 2D and 3D. Our method consistently outperforms a series of baselines on challenging benchmarks, such as retinal or whole-brain vessel graph extraction.
翻译:直接实现图像到图的变换是一项极具挑战的任务,它要求在单一模型中同时完成目标检测与关系预测。由于该任务的复杂性,许多领域内缺乏大规模训练数据集,使得大型网络的训练面临困难。数据稀疏性问题亟需建立类似于计算机视觉领域前沿的预训练策略。本文提出了一套能够实现图像到图变换器跨域与跨维度迁移学习的方法。我们提出了:(1) 一种正则化边采样损失函数,用于跨域采样最优数量的对象关系(边);(2) 一个面向图像到图变换器的域适应框架,能够对齐不同域的特征;(3) 一个简单的投影函数,允许我们在二维输入数据上预训练三维变换器。我们通过跨域与跨维度的实验验证了该方法的效果:在二维卫星图像上预训练模型后,将其应用于截然不同的二维与三维目标域。该方法在视网膜血管图提取、全脑血管图提取等挑战性基准任务中,始终优于多种基线方法。