Recent research suggests that combining AI models with a human expert can exceed the performance of either alone. The combination of their capabilities is often realized by learning to defer algorithms that enable the AI to learn to decide whether to make a prediction for a particular instance or defer it to the human expert. However, to accurately learn which instances should be deferred to the human expert, a large number of expert predictions that accurately reflect the expert's capabilities are required -- in addition to the ground truth labels needed to train the AI. This requirement shared by many learning to defer algorithms hinders their adoption in scenarios where the responsible expert regularly changes or where acquiring a sufficient number of expert predictions is costly. In this paper, we propose a three-step approach to reduce the number of expert predictions required to train learning to defer algorithms. It encompasses (1) the training of an embedding model with ground truth labels to generate feature representations that serve as a basis for (2) the training of an expertise predictor model to approximate the expert's capabilities. (3) The expertise predictor generates artificial expert predictions for instances not yet labeled by the expert, which are required by the learning to defer algorithms. We evaluate our approach on two public datasets. One with "synthetically" generated human experts and another from the medical domain containing real-world radiologists' predictions. Our experiments show that the approach allows the training of various learning to defer algorithms with a minimal number of human expert predictions. Furthermore, we demonstrate that even a small number of expert predictions per class is sufficient for these algorithms to exceed the performance the AI and the human expert can achieve individually.
翻译:近期研究表明,将AI模型与人类专家相结合可超越两者单独运行的性能。这种能力协同通常通过"学习延迟决策"算法实现——该算法使AI能够自主学习是否应对特定实例进行预测,或将其延迟至人类专家处理。然而,为了准确学习哪些实例应延迟至人类专家,除了训练AI所需的地面真值标签外,还需大量能精确反映专家能力的专家预测数据。这一多数学习延迟算法共有的需求,阻碍了其在责任专家频繁变更或获取足量专家预测成本高昂场景中的应用。本文提出一种三阶段方法来减少训练学习延迟算法所需的专家预测数量,包括:(1) 使用地面真值标签训练嵌入模型生成特征表示,为(2) 训练专家能力预测模型以近似专家能力提供基础;(3) 专家预测模型为专家尚未标注的实例生成人工专家预测,以满足学习延迟算法的需求。我们在两个公开数据集上评估本方法:一个包含"合成"生成的人类专家,另一个来自医疗领域并包含真实放射科医师的预测。实验表明,本方法能以最少的人类专家预测数量训练各类学习延迟算法。此外,我们证明即使每个类别仅需少量专家预测,这些算法也能超越AI与人类专家各自独立所能达到的性能。