We argue that many general evaluation problems can be viewed through the lens of voting theory. Each task is interpreted as a separate voter, which requires only ordinal rankings or pairwise comparisons of agents to produce an overall evaluation. By viewing the aggregator as a social welfare function, we are able to leverage centuries of research in social choice theory to derive principled evaluation frameworks with axiomatic foundations. These evaluations are interpretable and flexible, while avoiding many of the problems currently facing cross-task evaluation. We apply this Voting-as-Evaluation (VasE) framework across multiple settings, including reinforcement learning, large language models, and humans. In practice, we observe that VasE can be more robust than popular evaluation frameworks (Elo and Nash averaging), discovers properties in the evaluation data not evident from scores alone, and can predict outcomes better than Elo in a complex seven-player game. We identify one particular approach, maximal lotteries, that satisfies important consistency properties relevant to evaluation, is computationally efficient (polynomial in the size of the evaluation data), and identifies game-theoretic cycles
翻译:我们论证许多通用评估问题可通过投票理论的视角加以审视。每个任务被视作独立选民,仅需对智能体进行序数排序或两两比较即可生成整体评估。通过将聚合器视为社会福利函数,我们得以借鉴数百年社会选择理论研究成果,构建具有公理化基础的规范性评估框架。该评估兼具可解释性与灵活性,同时规避了当前跨任务评估面临的诸多问题。我们将这种"投票即评估"(VasE)框架应用于强化学习、大语言模型及人类行为等多个场景。实践表明,VasE框架比主流评估方法(Elo评分与纳什平均)更具鲁棒性,可发现评分数据中隐现的评估属性,并在复杂七玩家博弈中展现出优于Elo评分的预测能力。我们特别识别出最大抽签法这一典型方法,它满足评估相关的关键一致性属性,具有计算高效性(评估数据规模的多项式复杂度),并能识别博弈论中的循环现象。