Contrastive pretraining is well-known to improve downstream task performance and model generalisation, especially in limited label settings. However, it is sensitive to the choice of augmentation pipeline. Positive pairs should preserve semantic information while destroying domain-specific information. Standard augmentation pipelines emulate domain-specific changes with pre-defined photometric transformations, but what if we could simulate realistic domain changes instead? In this work, we show how to utilise recent progress in counterfactual image generation to this effect. We propose CF-SimCLR, a counterfactual contrastive learning approach which leverages approximate counterfactual inference for positive pair creation. Comprehensive evaluation across five datasets, on chest radiography and mammography, demonstrates that CF-SimCLR substantially improves robustness to acquisition shift with higher downstream performance on both in- and out-of-distribution data, particularly for domains which are under-represented during training.
翻译:对比预训练因能够提升下游任务性能与模型泛化能力而广为人知,尤其在标签受限场景下表现突出。然而,该方法对数据增强管线的选择高度敏感:正样本对需在保留语义信息的同时破坏领域特定特征。标准增强管线通过预定义光度变换模拟领域特异性变化,但若取而代之模拟真实领域变化会怎样?本研究展示了如何利用反事实图像生成的最新进展实现此目标。我们提出CF-SimCLR——一种利用近似反事实推理构建正样本对的反事实对比学习方法。基于胸部X光与乳腺X线摄影五种数据集的综合评估表明,CF-SimCLR显著提升了对采集偏移的稳健性,在分布内与分布外数据上均取得更优下游性能,尤其适用于训练中代表性不足的领域。