With the recent advancements of technology in facilitating real-time monitoring and data collection, "just-in-time" interventions can be delivered via mobile devices to achieve both real-time and long-term management and control. Reinforcement learning formalizes such mobile interventions as a sequence of decision rules and assigns treatment arms based on the user's status at each decision point. In practice, real applications concern a large number of decision points beyond the time horizon of the currently collected data. This usually refers to reinforcement learning in the infinite horizon setting, which becomes much more challenging. This article provides a selective overview of some statistical methodologies on this topic. We discuss their modeling framework, generalizability, and interpretability and provide some use case examples. Some future research directions are discussed in the end.
翻译:随着实时监测和数据收集技术的近期发展,“即时”干预可通过移动设备实现,从而达成实时与长期的管理和控制。强化学习将此类移动干预形式化为一组决策规则序列,并根据用户在每次决策点的状态分配治疗组别。在实际应用中,真实场景涉及超出当前收集数据时间域的大量决策点,这通常对应无限时间域设置下的强化学习,其挑战性显著增加。本文选择性概述了该主题的若干统计方法论,探讨了其建模框架、泛化能力与可解释性,并提供了若干应用案例。最后讨论了未来的研究方向。