Asymmetric reinforcement learning leverages privileged information available during training to improve learning under partial observability. Existing asymmetric actor-critic methods typically assume access to the full environment state to condition the critic during training, which is often unrealistic in practice. We introduce the informed asymmetric actor-critic framework that allows the critic to be conditioned on arbitrary state-dependent privileged signals, and show that any such signal yields unbiased policy gradient estimates. This substantially expands the set of admissible privileged information and raises the problem of selecting the most informative signals for learning. To this end, we propose two novel informativeness criteria: a dependence-based test that can be applied prior to training, and a test based on improvements in value prediction that can be applied post hoc. Experiments on partially observable benchmarks and synthetic environments demonstrate that carefully selected privileged signals can match or outperform full-state asymmetric baselines while relying on strictly less state information.
翻译:非对称强化学习利用训练阶段可获得的特权信息来改善部分可观测条件下的学习效果。现有非对称演员-评论家方法通常假设训练时可以获取完整环境状态来调节评论家,这在实际应用中往往不切实际。我们提出知情非对称演员-评论家框架,允许评论家基于任意状态依赖的特权信号进行调节,并证明任何此类信号都能产生无偏策略梯度估计。这显著扩展了可采纳特权信息集合,同时引出了如何选择最具信息量的信号进行学习的问题。为此,我们提出两种新型信息量评估准则:可应用于训练前的基于依赖性的检验方法,以及可应用于训练后的基于值预测改进的检验方法。在部分可观测基准测试及合成环境中的实验表明,精心选择的特权信号在使用严格更少状态信息的前提下,能够达到甚至超越全状态非对称基线方法的性能。