In reinforcement learning and imitation learning, an object of central importance is the state distribution induced by the policy. It plays a crucial role in the policy gradient theorem, and references to it--along with the related state-action distribution--can be found all across the literature. Despite its importance, the state distribution is mostly discussed indirectly and theoretically, rather than being modeled explicitly. The reason being an absence of appropriate density estimation tools. In this work, we investigate applications of a normalizing flow-based model for the aforementioned distributions. In particular, we use a pair of flows coupled through the optimality point of the Donsker-Varadhan representation of the Kullback-Leibler (KL) divergence, for distribution matching based imitation learning. Our algorithm, Coupled Flow Imitation Learning (CFIL), achieves state-of-the-art performance on benchmark tasks with a single expert trajectory and extends naturally to a variety of other settings, including the subsampled and state-only regimes.
翻译:在强化学习和模仿学习中,一个核心关键对象是由策略诱导的状态分布。它在策略梯度定理中起着关键作用,并且相关文献中广泛提及该分布及其对应的状态-动作分布。尽管其重要性显著,但状态分布大多仅被间接讨论和理论分析,而非显式建模,原因在于缺乏合适的密度估计工具。本研究探讨了基于归一化流的模型在上述分布中的应用。具体而言,我们通过Donsker-Varadhan表示的Kullback-Leibler散度最优性点耦合一对流,用于基于分布匹配的模仿学习。所提出的算法——耦合流模仿学习(CFIL)——在仅使用单条专家轨迹的基准任务上达到了当前最优性能,并能自然扩展到包括子采样和仅状态设置在内的多种其他场景。