Reinforcement learning has been increasingly applied in monitoring applications because of its ability to learn from previous experiences and can make adaptive decisions. However, existing machine learning-based health monitoring applications are mostly supervised learning algorithms, trained on labels and they cannot make adaptive decisions in an uncertain complex environment. This study proposes a novel and generic system, predictive deep reinforcement learning (PDRL) with multiple RL agents in a time series forecasting environment. The proposed generic framework accommodates virtual Deep Q Network (DQN) agents to monitor predicted future states of a complex environment with a well-defined reward policy so that the agent learns existing knowledge while maximizing their rewards. In the evaluation process of the proposed framework, three DRL agents were deployed to monitor a subject's future heart rate, respiration, and temperature predicted using a BiLSTM model. With each iteration, the three agents were able to learn the associated patterns and their cumulative rewards gradually increased. It outperformed the baseline models for all three monitoring agents. The proposed PDRL framework is able to achieve state-of-the-art performance in the time series forecasting process. The proposed DRL agents and deep learning model in the PDRL framework are customized to implement the transfer learning in other forecasting applications like traffic and weather and monitor their states. The PDRL framework is able to learn the future states of the traffic and weather forecasting and the cumulative rewards are gradually increasing over each episode.
翻译:强化学习因其能从过往经验中学习并做出自适应决策的能力,在监测应用中日益普及。然而,现有基于机器学习的健康监测应用大多采用监督学习算法,依赖标签进行训练,无法在不确定的复杂环境中做出自适应决策。本研究提出一种新颖且通用的系统——预测性深度强化学习(PDRL),其在时间序列预测环境中部署多个强化学习智能体。该通用框架容纳虚拟深度Q网络(DQN)智能体,通过定义明确的奖励策略来监测复杂环境的预测未来状态,使智能体在学习既有知识的同时最大化其奖励。在评估该框架时,三个深度强化学习(DRL)智能体被部署用于监测受试者未来心率、呼吸和体温的预测值(采用双向长短期记忆网络BiLSTM模型)。每次迭代中,三个智能体均能学习关联模式,其累积奖励逐步增加。该框架在所有三个监测任务中均优于基线模型。所提出的PDRL框架在时间序列预测过程中能够达到当前最优性能。PDRL框架中的DRL智能体及深度学习模型可根据需要定制,以在交通、天气等其他预测应用中实现迁移学习并监测其状态。PDRL框架能够学习交通和天气预测的未来状态,其累积奖励随每次回合逐步增加。