Off-Policy Evaluation (OPE) aims to assess the effectiveness of counterfactual policies using only offline logged data and is often used to identify the top-k promising policies for deployment in online A/B tests. Existing evaluation metrics for OPE estimators primarily focus on the "accuracy" of OPE or that of downstream policy selection, neglecting risk-return tradeoff in the subsequent online policy deployment. To address this issue, we draw inspiration from portfolio evaluation in finance and develop a new metric, called SharpeRatio@k, which measures the risk-return tradeoff of policy portfolios formed by an OPE estimator under varying online evaluation budgets (k). We validate our metric in two example scenarios, demonstrating its ability to effectively distinguish between low-risk and high-risk estimators and to accurately identify the most efficient estimator. This efficient estimator is characterized by its capability to form the most advantageous policy portfolios, maximizing returns while minimizing risks during online deployment, a nuance that existing metrics typically overlook. To facilitate a quick, accurate, and consistent evaluation of OPE via SharpeRatio@k, we have also integrated this metric into an open-source software, SCOPE-RL. Employing SharpeRatio@k and SCOPE-RL, we conduct comprehensive benchmarking experiments on various estimators and RL tasks, focusing on their risk-return tradeoff. These experiments offer several interesting directions and suggestions for future OPE research.
翻译:离线策略评估旨在利用仅有的离线日志数据评估反事实策略的有效性,通常用于识别排名前k的最有前景策略以部署至在线A/B测试。现有离线策略评估估计器的评价指标主要关注OPE本身的"准确性"或下游策略选择的准确性,忽略了后续在线策略部署中的风险-收益权衡。为解决该问题,我们从金融领域的投资组合评估中汲取灵感,提出一种名为SharpeRatio@k的新指标,用于衡量在不同在线评估预算(k)下由OPE估计器形成的策略投资组合的风险-收益权衡。我们在两个示例场景中验证了该指标,证明其能有效区分低风险与高风险估计器,并准确识别最优效率估计器。该高效估计器的特性在于能够形成最具优势的策略投资组合,在在线部署中实现收益最大化同时最小化风险——这一细微差异在现有指标中通常被忽略。为通过SharpeRatio@k快速、准确且一致地评估OPE,我们已将该指标集成至开源软件SCOPE-RL中。通过运用SharpeRatio@k与SCOPE-RL,我们针对多种估计器与强化学习任务开展了全面的基准测试实验,聚焦其风险-收益权衡。这些实验为未来OPE研究提供了若干有趣的方向与建议。