We consider a simulation optimization problem for a context-dependent decision-making, which aims to determine the top-m designs for all contexts. Under a Bayesian framework, we formulate the optimal dynamic sampling decision as a stochastic dynamic programming problem, and develop a sequential sampling policy to efficiently learn the performance of each design under each context. The asymptotically optimal sampling ratios are derived to attain the optimal large deviations rate of the worst-case of probability of false selection. The proposed sampling policy is proved to be consistent and its asymptotic sampling ratios are asymptotically optimal. Numerical experiments demonstrate that the proposed method improves the efficiency for selection of top-m context-dependent designs.
翻译:我们考虑一个面向上下文相关决策的仿真优化问题,其目标是在所有上下文中确定前m个最优设计。在贝叶斯框架下,我们将最优动态采样决策建模为随机动态规划问题,并开发了一种序贯采样策略,以高效学习每个设计在每个上下文下的性能。通过推导渐近最优采样比,我们实现了最差情况下错误选择概率的渐近最优大偏差率。所提出的采样策略被证明具有一致性,且其渐近采样比具有渐近最优性。数值实验表明,该方法提高了选择Top-m上下文相关设计的效率。