This thesis considers sequential decision problems, where the loss/reward incurred by selecting an action may not be inferred from observed feedback. A major part of this thesis focuses on the unsupervised sequential selection problem, where one can not infer the loss incurred for selecting an action from observed feedback. We also introduce a new setup named Censored Semi Bandits, where the loss incurred for selecting an action can be observed under certain conditions. Finally, we study the channel selection problem in the communication networks, where the reward for an action is only observed when no other player selects that action to play in the round. These problems find applications in many fields like healthcare, crowd-sourcing, security, adaptive resource allocation, among many others. This thesis aims to address the above-described sequential decision problems by exploiting specific structures these problems exhibit. We develop provably optimal algorithms for each of these setups with weak feedback and validate their empirical performance on different problem instances derived from synthetic and real datasets.
翻译:本文研究序贯决策问题,在该问题中,选择某个动作所产生的损失/收益可能无法从观测到的反馈中推断得出。论文的主要部分聚焦于无监督序贯选择问题,即无法从观测反馈中推断出所选动作的损失。此外,我们引入了一个名为“删失式半强盗”的新设定,在该设定下,选择某个动作所产生的损失可以在特定条件下被观测到。最后,我们研究了通信网络中的信道选择问题,在该问题中,只有当回合中没有其他玩家选择该动作时,才能观测到该动作的收益。这些问题在医疗健康、众包、安全、自适应资源分配等多个领域均有应用。本文旨在通过利用上述问题所具备的特定结构,来解决这些序贯决策问题。我们针对每种具有弱反馈的设定开发了可证明最优的算法,并在基于合成数据集和真实数据集生成的多个不同问题实例上验证了其实验性能。