Over the recent past data-driven algorithms for solving stochastic optimal control problems in face of model uncertainty have become an increasingly active area of research. However, for singular controls and underlying diffusion dynamics the analysis has so far been restricted to the scalar case. In this paper we fill this gap by studying a multivariate singular control problem for reversible diffusions with controls of reflection type. Our contributions are threefold. We first explicitly determine the long-run average costs as a domain-dependent functional, showing that the control problem can be equivalently characterized as a shape optimization problem. For given diffusion dynamics, assuming the optimal domain to be strongly star-shaped, we then propose a gradient descent algorithm based on polytope approximations to numerically determine a cost-minimizing domain. Finally, we investigate data-driven solutions when the diffusion dynamics are unknown to the controller. Using techniques from nonparametric statistics for stochastic processes, we construct an optimal domain estimator, whose static regret is bounded by the minimax optimal estimation rate of the unreflected process' invariant density. In the most challenging situation, when the dynamics must be learned simultaneously to controlling the process, we develop an episodic learning algorithm to overcome the emerging exploration-exploitation dilemma and show that given the static regret as a baseline, the loss in its sublinear regret per time unit is of natural order compared to the one-dimensional case.
翻译:近年来,在模型不确定性下求解随机最优控制问题的数据驱动算法已成为一个日益活跃的研究领域。然而,对于奇异控制及底层扩散动力学,此前相关分析仅局限于标量情形。本文通过研究一类反射型控制的可逆扩散多变量奇异控制问题填补了这一空白。我们的贡献体现在三个方面:首先,我们明确地将长期平均成本表示为域依赖泛函,证明该控制问题可等价转化为形状优化问题;其次,针对给定扩散动力学且假设最优域为强星形域的情形,我们提出一种基于多面体近似的梯度下降算法以数值求解成本最小化域;最后,我们研究了控制器未知扩散动力学时的数据驱动解。利用随机过程的非参数统计技术,我们构建了最优域估计器,其静态遗憾被无反射过程不变密度的极小最优估计率所界定。在最具挑战性的情形——即需在控制过程的同时学习动力学时,我们开发了一种情节式学习算法以克服探索-利用困境,并证明以静态遗憾为基准,其每时间单位次线性遗憾的损失相较于单维情形具有自然阶次。