We introduce a novel framework for incorporating human expertise into algorithmic predictions. Our approach focuses on the use of human judgment to distinguish inputs which `look the same' to any feasible predictive algorithm. We argue that this framing clarifies the problem of human/AI collaboration in prediction tasks, as experts often have access to information -- particularly subjective information -- which is not encoded in the algorithm's training data. We use this insight to develop a set of principled algorithms for selectively incorporating human feedback only when it improves the performance of any feasible predictor. We find empirically that although algorithms often outperform their human counterparts on average, human judgment can significantly improve algorithmic predictions on specific instances (which can be identified ex-ante). In an X-ray classification task, we find that this subset constitutes nearly 30% of the patient population. Our approach provides a natural way of uncovering this heterogeneity and thus enabling effective human-AI collaboration.
翻译:本文提出了一种将人类专业知识融入算法预测的新框架。我们的方法侧重于利用人类判断来区分那些对于任何可行的预测算法都"看起来相同"的输入。我们认为,这种框架能够澄清预测任务中人机协作的问题,因为专家通常能够获取算法训练数据中未编码的信息——特别是主观信息。基于这一洞见,我们开发了一套原则性算法,仅在人类反馈能够提升任何可行预测器性能时,才选择性地将其纳入。实证研究表明,虽然算法在平均表现上通常优于人类,但在特定实例(可事先识别)上,人类判断能够显著提升算法预测的准确性。在X射线分类任务中,我们发现这类实例约占患者群体的30%。我们的方法为揭示这种异质性提供了一种自然途径,从而实现了有效的人机协作。