In rank aggregation (RA), a collection of preferences from different users are summarized into a total order under the assumption of homogeneity of users. Model misspecification in RA arises since the homogeneity assumption fails to be satisfied in the complex real-world situation. Existing robust RAs usually resort to an augmentation of the ranking model to account for additional noises, where the collected preferences can be treated as a noisy perturbation of idealized preferences. Since the majority of robust RAs rely on certain perturbation assumptions, they cannot generalize well to agnostic noise-corrupted preferences in the real world. In this paper, we propose CoarsenRank, which possesses robustness against model misspecification. Specifically, the properties of our CoarsenRank are summarized as follows: (1) CoarsenRank is designed for mild model misspecification, which assumes there exist the ideal preferences (consistent with model assumption) that locates in a neighborhood of the actual preferences. (2) CoarsenRank then performs regular RAs over a neighborhood of the preferences instead of the original dataset directly. Therefore, CoarsenRank enjoys robustness against model misspecification within a neighborhood. (3) The neighborhood of the dataset is defined via their empirical data distributions. Further, we put an exponential prior on the unknown size of the neighborhood, and derive a much-simplified posterior formula for CoarsenRank under particular divergence measures. (4) CoarsenRank is further instantiated to Coarsened Thurstone, Coarsened Bradly-Terry, and Coarsened Plackett-Luce with three popular probability ranking models. Meanwhile, tractable optimization strategies are introduced with regards to each instantiation respectively. In the end, we apply CoarsenRank on four real-world datasets.
翻译:在排序聚合中,不同用户的偏好集合在用户同质性假设下被汇总为一个全序。然而,由于复杂现实情境中同质性假设无法满足,排序聚合中会出现模型误设问题。现有鲁棒排序聚合方法通常通过扩展排序模型来引入额外噪声,将收集到的偏好视为理想偏好的噪声扰动。由于多数鲁棒方法依赖于特定扰动假设,因此难以泛化至现实中不可知的噪声污染偏好。本文提出CoarsenRank方法,其具备对模型误设的鲁棒性。具体而言,CoarsenRank的特性可总结如下:(1)CoarsenRank针对轻度模型误设设计,假设存在与模型假设一致的理想偏好,且该理想偏好位于实际偏好的邻域内。(2)CoarsenRank在偏好邻域而非原始数据集上直接执行常规排序聚合,从而在邻域范围内抵御模型误设的影响。(3)数据集的邻域通过其经验数据分布定义。进一步,我们对邻域未知大小施加指数先验,并在特定散度度量下推导出CoarsenRank的高度简化后验公式。(4)CoarsenRank被具体实例化为Coarsened Thurstone、Coarsened Bradley-Terry和Coarsened Plackett-Luce三种基于流行概率排序模型的方法,并针对每种实例化引入可操作的优化策略。最后,我们在四个真实数据集上应用CoarsenRank进行验证。