Many real-world applications of reinforcement learning (RL) require making decisions in continuous action environments. In particular, determining the optimal dose level plays a vital role in developing medical treatment regimes. One challenge in adapting existing RL algorithms to medical applications, however, is that the popular infinite support stochastic policies, e.g., Gaussian policy, may assign riskily high dosages and harm patients seriously. Hence, it is important to induce a policy class whose support only contains near-optimal actions, and shrink the action-searching area for effectiveness and reliability. To achieve this, we develop a novel \emph{quasi-optimal learning algorithm}, which can be easily optimized in off-policy settings with guaranteed convergence under general function approximations. Theoretically, we analyze the consistency, sample complexity, adaptability, and convergence of the proposed algorithm. We evaluate our algorithm with comprehensive simulated experiments and a dose suggestion real application to Ohio Type 1 diabetes dataset.
翻译:强化学习(RL)的许多实际应用需要在连续动作环境中进行决策。特别是,确定最佳剂量水平在制定医疗治疗方案中起着至关重要的作用。然而,将现有强化学习算法应用于医疗领域面临一个挑战:流行的具有无限支撑的随机策略(例如高斯策略)可能会分配存在风险的过高剂量,从而严重危害患者。因此,有必要引入一种支撑仅包含近最优动作的策略类,并缩小动作搜索范围以提高有效性和可靠性。为此,我们提出了一种新颖的\emph{准最优学习算法},该算法可在离策略设置下轻松优化,并在通用函数近似下保证收敛。理论上,我们分析了所提算法的一致性、样本复杂度、适应性和收敛性。通过综合仿真实验以及针对俄亥俄1型糖尿病数据集的真实剂量建议应用,我们对算法进行了评估。