The partial monitoring (PM) framework provides a theoretical formulation of sequential learning problems with incomplete feedback. On each round, a learning agent plays an action while the environment simultaneously chooses an outcome. The agent then observes a feedback signal that is only partially informative about the (unobserved) outcome. The agent leverages the received feedback signals to select actions that minimize the (unobserved) cumulative loss. In contextual PM, the outcomes depend on some side information that is observable by the agent before selecting the action on each round. In this paper, we consider the contextual and non-contextual PM settings with stochastic outcomes. We introduce a new class of strategies based on the randomization of deterministic confidence bounds, that extend regret guarantees to settings where existing stochastic strategies are not applicable. Our experiments show that the proposed RandCBP and RandCBPside* strategies improve state-of-the-art baselines in PM games. To encourage the adoption of the PM framework, we design a use case on the real-world problem of monitoring the error rate of any deployed classification system.
翻译:部分监控(PM)框架为不完全反馈下的序贯学习问题提供了理论形式化描述。在每个时间步,学习智能体执行一个动作,而环境同时选择一种结果。智能体随后观察到仅部分揭示(未观测)结果的反馈信号。智能体利用接收到的反馈信号选择能最小化(未观测)累积损失的动作。在上下文PM中,结果依赖于智能体在每个时间步选择动作前可观测的辅助信息。本文研究了具有随机结果的上下文与非上下文PM设定。我们提出了一类基于确定性置信界随机化的新策略,将遗憾保证扩展至现有随机策略无法适用的场景。实验表明,所提出的RandCBP与RandCBPside*策略在PM博弈中改进了现有最优基线方法。为促进PM框架的应用,我们设计了一个针对监控任一部署分类系统错误率的真实世界问题的使用案例。