We present a novel method aimed at enhancing the sample efficiency of ensemble Q learning. Our proposed approach integrates multi-head self-attention into the ensembled Q networks while bootstrapping the state-action pairs ingested by the ensemble. This not only results in performance improvements over the original REDQ (Chen et al. 2021) and its variant DroQ (Hi-raoka et al. 2022), thereby enhancing Q predictions, but also effectively reduces both the average normalized bias and standard deviation of normalized bias within Q-function ensembles. Importantly, our method also performs well even in scenarios with a low update-to-data (UTD) ratio. Notably, the implementation of our proposed method is straightforward, requiring minimal modifications to the base model.
翻译:我们提出一种旨在提升集成Q学习样本效率的新方法。该方法将多头自注意力机制引入集成Q网络,同时对集成模型所摄入的状态-动作对实施自助法采样。这不仅在原始REDQ(Chen等,2021)及其变体DroQ(Hiraoka等,2022)基础上实现了性能提升,从而优化Q值预测,还能有效降低Q函数集成内的平均归一化偏差及归一化偏差标准差。重要的是,即使在低更新-数据比(UTD)场景下,本方法仍表现优异。值得注意的是,所提方法实现简洁,仅需对基础模型进行最小化修改。