Public dataset limitations have significantly hindered the development and benchmarking of learning to defer (L2D) algorithms, which aim to optimally combine human and AI capabilities in hybrid decision-making systems. In such systems, human availability and domain-specific concerns introduce difficulties, while obtaining human predictions for training and evaluation is costly. Financial fraud detection is a high-stakes setting where algorithms and human experts often work in tandem; however, there are no publicly available datasets for L2D concerning this important application of human-AI teaming. To fill this gap in L2D research, we introduce the Financial Fraud Alert Review Dataset (FiFAR), a synthetic bank account fraud detection dataset, containing the predictions of a team of 50 highly complex and varied synthetic fraud analysts, with varied bias and feature dependence. We also provide a realistic definition of human work capacity constraints, an aspect of L2D systems that is often overlooked, allowing for extensive testing of assignment systems under real-world conditions. We use our dataset to develop a capacity-aware L2D method and rejection learning approach under realistic data availability conditions, and benchmark these baselines under an array of 300 distinct testing scenarios. We believe that this dataset will serve as a pivotal instrument in facilitating a systematic, rigorous, reproducible, and transparent evaluation and comparison of L2D methods, thereby fostering the development of more synergistic human-AI collaboration in decision-making systems. The public dataset and detailed synthetic expert information are available at: https://github.com/feedzai/fifar-dataset
翻译:公共数据集的局限性严重阻碍了学习延迟决策(L2D)算法的开发与基准测试,该类算法旨在混合式人机决策系统中实现人类与AI能力的最优协同。在此类系统中,人类参与者的可用性及领域特定问题带来诸多挑战,而为训练和评估获取人工预测的成本高昂。金融欺诈检测是一个高风险场景,算法与人类专家常需协同工作,但针对这一人机协作重要应用领域,目前尚无公开可用的L2D数据集。为填补L2D研究领域的这一空白,我们提出了金融欺诈预警复核数据集(FiFAR),这是一个合成银行账户欺诈检测数据集,包含由50名具有高度复杂性和多样性的合成欺诈分析师组成的团队所生成的预测结果,这些分析师具有不同的偏见和特征依赖关系。我们还提供了人类工作容量约束的现实定义——这是L2D系统中常被忽视的一个方面,从而允许在真实条件下对分配系统进行广泛测试。我们利用该数据集,在真实数据可用性条件下开发了容量感知型L2D方法及拒绝学习方案,并在300种不同测试场景下对这些基准方法进行了评估。我们相信,该数据集将作为关键工具,促进对L2D方法进行系统、严格、可重复且透明的评估与比较,进而推动决策系统中更具协同性的人机协作发展。公开数据集及详细的合成专家信息获取地址:https://github.com/feedzai/fifar-dataset