Reinforcement learning (RL) problems where the learner attempts to infer an unobserved reward from some feedback variables have been studied in several recent papers. The setting of Interaction-Grounded Learning (IGL) is an example of such feedback-based reinforcement learning tasks where the learner optimizes the return by inferring latent binary rewards from the interaction with the environment. In the IGL setting, a relevant assumption used in the RL literature is that the feedback variable $Y$ is conditionally independent of the context-action $(X,A)$ given the latent reward $R$. In this work, we propose Variational Information-based IGL (VI-IGL) as an information-theoretic method to enforce the conditional independence assumption in the IGL-based RL problem. The VI-IGL framework learns a reward decoder using an information-based objective based on the conditional mutual information (MI) between the context-action $(X,A)$ and the feedback variable $Y$ observed from the environment. To estimate and optimize the information-based terms for the continuous random variables in the RL problem, VI-IGL leverages the variational representation of mutual information and results in a min-max optimization problem. Furthermore, we extend the VI-IGL framework to general $f$-Information measures in the information theory literature, leading to the generalized $f$-VI-IGL framework to address the RL problem under the IGL condition. Finally, we provide the empirical results of applying the VI-IGL method to several reinforcement learning settings, which indicate an improved performance in comparison to the previous IGL-based RL algorithm.
翻译:在近期的多项研究中,研究者们探索了强化学习(RL)中学习器从某些反馈变量中推断未观测奖励的问题。交互式强化学习(IGL)设定正是这类基于反馈的强化学习任务的一个实例——学习器通过与环境的交互推断潜在二元奖励,从而优化累积回报。在IGL设定中,反馈变量$Y$在给定潜在奖励$R$的条件下与上下文-动作$(X,A)$条件独立,这是强化学习文献中采用的关键假设。本文提出基于变分信息论的IGL方法(VI-IGL),这是一种通过信息论手段强化IGL型RL问题中条件独立性假设的方法。VI-IGL框架利用环境反馈变量$Y$与上下文-动作$(X,A)$之间的条件互信息(MI)构建信息论目标函数,以此学习奖励解码器。为估计并优化RL问题中连续随机变量的信息论项,VI-IGL采用互信息的变分表示,将其转化为极小极大优化问题。进一步地,我们将VI-IGL框架推广至信息论中通用的$f$-信息度量,构建了广义$f$-VI-IGL框架以解决IGL条件下的RL问题。最后,通过在多个强化学习设定中应用VI-IGL方法,实验结果表明该方法相较于现有IGL型RL算法具有更优性能。