We consider the problem of sequential evaluation, in which an evaluator observes candidates in a sequence and assigns scores to these candidates in an online, irrevocable fashion. Motivated by the psychology literature that has studied sequential bias in such settings -- namely, dependencies between the evaluation outcome and the order in which the candidates appear -- we propose a natural model for the evaluator's rating process that captures the lack of calibration inherent to such a task. We conduct crowdsourcing experiments to demonstrate various facets of our model. We then proceed to study how to correct sequential bias under our model by posing this as a statistical inference problem. We propose a near-linear time, online algorithm for this task and prove guarantees in terms of two canonical ranking metrics. We also prove that our algorithm is information theoretically optimal, by establishing matching lower bounds in both metrics. Finally, we perform a host of numerical experiments to show that our algorithm often outperforms the de facto method of using the rankings induced by the reported scores, both in simulation and on the crowdsourcing data that we collected.
翻译:我们考虑序列化评估问题,其中评估者按顺序观察候选人,并以在线且不可撤销的方式为这些候选人分配分数。受研究此类情境中序列偏差(即评估结果与候选人出现顺序之间的依赖关系)的心理学文献启发,我们提出了一个评估者评分过程的自然模型,该模型捕捉了此类任务固有的校准缺失问题。我们通过众包实验展示了模型的多个方面。随后,我们研究如何在该模型下修正序列偏差,将其视为一个统计推断问题。我们提出了一种近线性时间的在线算法,并从两个经典排序指标的角度证明了其性能保证。我们还通过为这两个指标建立匹配的下界,证明了该算法在信息论意义上的最优性。最后,我们进行了大量数值实验,展示我们的算法在模拟实验以及我们收集的众包数据上,通常优于使用报告分数所导出排序的常规方法。