Decision-making under uncertainty is a fundamental problem encountered frequently and can be formulated as a stochastic multi-armed bandit problem. In the problem, the learner interacts with an environment by choosing an action at each round, where a round is an instance of an interaction. In response, the environment reveals a reward, which is sampled from a stochastic process, to the learner. The goal of the learner is to maximize cumulative reward. In this work, we assume that the rewards are the inner product of an action vector and a state vector generated by a linear Gaussian dynamical system. To predict the reward for each action, we propose a method that takes a linear combination of previously observed rewards for predicting each action's next reward. We show that, regardless of the sequence of previous actions chosen, the reward sampled for any previously chosen action can be used for predicting another action's future reward, i.e. the reward sampled for action 1 at round $t-1$ can be used for predicting the reward for action $2$ at round $t$. This is accomplished by designing a modified Kalman filter with a matrix representation that can be learned for reward prediction. Numerical evaluations are carried out on a set of linear Gaussian dynamical systems and are compared with 2 other well-known stochastic multi-armed bandit algorithms.
翻译:不确定性下的决策制定是一个频繁遇到的基本问题,可被表述为随机多臂老虎机问题。在该问题中,学习者通过每轮选择一个动作与环境进行交互,其中一轮即一次交互实例。作为响应,环境会向学习者展示一个从随机过程中采样的奖励。学习者的目标是最大化累积奖励。在本研究中,我们假设奖励是动作向量与由线性高斯动态系统生成的状态向量的内积。为预测每个动作的奖励,我们提出一种方法,该方法采用先前观测到的奖励的线性组合来预测每个动作的下一次奖励。我们证明,无论先前选择的动作序列如何,为任何先前选择的动作采样的奖励都可用于预测另一动作的未来奖励,即第 $t-1$ 轮为动作 $1$ 采样的奖励可用于预测第 $t$ 轮动作 $2$ 的奖励。这是通过设计一个具有可学习矩阵表示的改进卡尔曼滤波器来实现的,该滤波器可用于奖励预测。我们在多组线性高斯动态系统上进行了数值评估,并与另外两种著名的随机多臂老虎机算法进行了比较。