Existing episodic reinforcement algorithms assume that the length of an episode is fixed across time and known a priori. In this paper, we consider a general framework of episodic reinforcement learning when the length of each episode is drawn from a distribution. We first establish that this problem is equivalent to online reinforcement learning with general discounting where the learner is trying to optimize the expected discounted sum of rewards over an infinite horizon, but where the discounting function is not necessarily geometric. We show that minimizing regret with this new general discounting is equivalent to minimizing regret with uncertain episode lengths. We then design a reinforcement learning algorithm that minimizes regret with general discounting but acts for the setting with uncertain episode lengths. We instantiate our general bound for different types of discounting, including geometric and polynomial discounting. We also show that we can obtain similar regret bounds even when the uncertainty over the episode lengths is unknown, by estimating the unknown distribution over time. Finally, we compare our learning algorithms with existing value-iteration based episodic RL algorithms in a grid-world environment.
翻译:现有基于片段的强化学习算法假设片段长度随时间固定且先验已知。本文考虑片段长度服从某种分布的一般性在线强化学习框架。首先,我们证明该问题等价于具有广义折扣因子的在线强化学习——学习器试图在无限时间范围内优化期望折扣奖励总和,但其折扣函数不一定是几何序列。研究表明,最小化这种新广义折扣机制下的遗憾值,等价于最小化片段长度不确定情况下的遗憾值。在此基础上,我们设计了一种能在广义折扣机制下最小化遗憾值、但针对片段长度不确定场景运行的强化学习算法。针对包括几何折扣和多项式折扣在内的不同折扣类型,我们给出了通用边界的具体实例。即便在片段长度不确定性未知的情况下,通过随时间估计未知分布,我们也能获得相似的遗憾值界限。最后,在网格世界环境中,我们将所提出的学习算法与现有基于值迭代的片段强化学习算法进行了比较。