Successful content moderation in online platforms relies on a human-AI collaboration approach. A typical heuristic estimates the expected harmfulness of a post and uses fixed thresholds to decide whether to remove it and whether to send it for human review. This disregards the prediction uncertainty, the time-varying element of human review capacity and post arrivals, and the selective sampling in the dataset (humans only review posts filtered by the admission algorithm). In this paper, we introduce a model to capture the human-AI interplay in content moderation. The algorithm observes contextual information for incoming posts, makes classification and admission decisions, and schedules posts for human review. Only admitted posts receive human reviews on their harmfulness. These reviews help educate the machine-learning algorithms but are delayed due to congestion in the human review system. The classical learning-theoretic way to capture this human-AI interplay is via the framework of learning to defer, where the algorithm has the option to defer a classification task to humans for a fixed cost and immediately receive feedback. Our model contributes to this literature by introducing congestion in the human review system. Moreover, unlike work on online learning with delayed feedback where the delay in the feedback is exogenous to the algorithm's decisions, the delay in our model is endogenous to both the admission and the scheduling decisions. We propose a near-optimal learning algorithm that carefully balances the classification loss from a selectively sampled dataset, the idiosyncratic loss of non-reviewed posts, and the delay loss of having congestion in the human review system. To the best of our knowledge, this is the first result for online learning in contextual queueing systems and hence our analytical framework may be of independent interest.
翻译:在线平台成功的内容审核依赖于人机协同方法。典型的启发式方法会评估帖子的预期危害性,并使用固定阈值来决定是否删除该帖子以及是否将其发送给人工审核。这种方法忽略了预测的不确定性、人工审核能力和帖子到达量的时变特性,以及数据集中的选择性抽样(人工仅审核经过准入算法筛选的帖子)。本文提出一个模型来刻画内容审核中的人机协同机制。该算法观察传入帖子的上下文信息,做出分类和准入决策,并安排帖子进行人工审核。只有被准入的帖子才会收到关于其危害性的人工审核。这些审核有助于指导机器学习算法,但由于人工审核系统的拥塞,审核结果会出现延迟。从经典学习理论角度刻画这种人机协同的框架是延迟学习,即算法可以选择以固定成本将分类任务延迟给人类处理,并立即获得反馈。我们的模型通过引入人工审核系统的拥塞机制对此研究领域作出贡献。此外,与反馈延迟外生于算法决策的在线延迟反馈学习研究不同,我们模型中的延迟内生于准入和调度决策。我们提出一种近似最优的学习算法,该算法仔细平衡了来自选择性抽样数据集的分类损失、未审核帖子的异质性损失以及人工审核系统拥塞造成的延迟损失。据我们所知,这是上下文排队系统中在线学习的首个研究成果,因此我们的分析框架可能具有独立的学术价值。