We consider a system of multiple sources, a single communication channel, and a single monitoring station. Each source measures a time-varying quantity with varying levels of accuracy and one of them sends its update to the monitoring station via the channel. The probability of success of each attempted communication is a function of the source scheduled for transmitting its update. Both the probability of correct measurement and the probability of successful transmission of all the sources are unknown to the scheduler. The metric of interest is the reward received by the system which depends on the accuracy of the last update received by the destination and the Age-of-Information (AoI) of the system. We model our scheduling problem as a variant of the multi-arm bandit problem with sources as different arms. We compare the performance of all $4$ standard bandit policies, namely, ETC, $\epsilon$-greedy, UCB, and TS suitably adjusted to our system model via simulations. In addition, we provide analytical guarantees of $2$ of these policies, ETC, and $\epsilon$-greedy. Finally, we characterize the lower bound on the cumulative regret achievable by any policy.
翻译:我们考虑一个包含多个数据源、单个通信信道和单个监测站的系统。每个数据源以不同精度测量时变量,其中一个数据源通过信道向监测站发送其更新信息。每次通信尝试的成功概率是调度发送更新的数据源的函数。对于所有数据源而言,正确测量的概率和成功传输的概率对调度器而言均未知。衡量指标是系统获得的奖励,该奖励取决于目的端接收到的最近更新的准确性和系统的信息时效(AoI)。我们将调度问题建模为以数据源为不同臂的多臂老虎机问题的变体。通过仿真,我们比较了所有四种标准老虎机策略(即ETC、ε-贪心、UCB和TS)在系统模型中的适配性能。此外,我们为其中两种策略(ETC和ε-贪心)提供了理论保障。最后,我们刻画了任意策略所能达到的累积遗憾下界。