Recommender Systems today are still mostly evaluated in terms of accuracy, with other aspects beyond the immediate relevance of recommendations, such as diversity, long-term user retention and fairness, often taking a back seat. Moreover, reconciling multiple performance perspectives is by definition indeterminate, presenting a stumbling block to those in the pursuit of rounded evaluation of Recommender Systems. EvalRS 2022 -- a data challenge designed around Multi-Objective Evaluation -- was a first practical endeavour, providing many insights into the requirements and challenges of balancing multiple objectives in evaluation. In this work, we reflect on EvalRS 2022 and expound upon crucial learnings to formulate a first-principles approach toward Multi-Objective model selection, and outline a set of guidelines for carrying out a Multi-Objective Evaluation challenge, with potential applicability to the problem of rounded evaluation of competing models in real-world deployments.
翻译:当前推荐系统的评估仍主要围绕准确性展开,而多样性、长期用户留存、公平性等超越即时相关性的维度常被置于次要地位。此外,协调多个性能视角在本质上存在不确定性,这为追求推荐系统全面评估的研究者设置了障碍。EvalRS 2022——一个围绕多目标评估设计的数据挑战赛——作为首次实践探索,为平衡评估中多目标的需求与挑战提供了丰富洞见。本文基于EvalRS 2022的反思,总结关键经验,提出基于第一性原理的多目标模型选择方法,并制定多目标评估挑战的实施指南。该指南对现实部署中竞争模型的全面评估问题具有潜在适用性。