Many real-time applications of the Internet of Things (IoT) need to deal with correlated information generated by multiple sensors. The design of efficient status update strategies that minimize the Age of Correlated Information (AoCI) is a key factor. In this paper, we consider an IoT network consisting of sensors equipped with the energy harvesting (EH) capability. We optimize the average AoCI at the data fusion center (DFC) by appropriately managing the energy harvested by sensors, whose true battery states are unobservable during the decision-making process. Particularly, we first formulate the dynamic status update procedure as a partially observable Markov decision process (POMDP), where the environmental dynamics are unknown to the DFC. In order to address the challenges arising from the causality of energy usage, unknown environmental dynamics, unobservability of sensors'true battery states, and large-scale discrete action space, we devise a deep reinforcement learning (DRL)-based dynamic status update algorithm. The algorithm leverages the advantages of the soft actor-critic and long short-term memory techniques. Meanwhile, it incorporates our proposed action decomposition and mapping mechanism. Extensive simulations are conducted to validate the effectiveness of our proposed algorithm by comparing it with available DRL algorithms for POMDPs.
翻译:许多物联网(IoT)实时应用需要处理由多个传感器生成的相关信息。设计最小化相关信息年龄(AoCI)的高效状态更新策略至关重要。本文考虑一个由具备能量采集(EH)能力的传感器构成的物联网网络。通过合理管理传感器所采集的能量(其真实电池状态在决策过程中不可观测),我们优化了数据融合中心(DFC)处的平均AoCI。具体而言,我们首先将动态状态更新过程建模为部分可观测马尔可夫决策过程(POMDP),其中环境动态对DFC未知。为应对能源使用的因果性、未知环境动态、传感器真实电池状态不可观测以及大规模离散动作空间带来的挑战,我们设计了一种基于深度强化学习(DRL)的动态状态更新算法。该算法融合了软演员-评论家(soft actor-critic)与长短期记忆技术的优势,同时整合了我们提出的动作分解与映射机制。通过将所提算法与现有用于POMDP的DRL算法进行对比,大量仿真实验验证了其有效性。