One of the most critical aspects of multimodal Reinforcement Learning (RL) is the effective integration of different observation modalities. Having robust and accurate representations derived from these modalities is key to enhancing the robustness and sample efficiency of RL algorithms. However, learning representations in RL settings for visuotactile data poses significant challenges, particularly due to the high dimensionality of the data and the complexity involved in correlating visual and tactile inputs with the dynamic environment and task objectives. To address these challenges, we propose Multimodal Contrastive Unsupervised Reinforcement Learning (M2CURL). Our approach employs a novel multimodal self-supervised learning technique that learns efficient representations and contributes to faster convergence of RL algorithms. Our method is agnostic to the RL algorithm, thus enabling its integration with any available RL algorithm. We evaluate M2CURL on the Tactile Gym 2 simulator and we show that it significantly enhances the learning efficiency in different manipulation tasks. This is evidenced by faster convergence rates and higher cumulative rewards per episode, compared to standard RL algorithms without our representation learning approach.
翻译:多模态强化学习中最关键的方面之一是有效整合不同的观测模态。从这些模态中获取鲁棒且准确的特征表示是提升强化学习算法鲁棒性和样本效率的关键。然而,在强化学习环境中为视觉触觉数据学习特征表示面临重大挑战,尤其是由于数据的高维性以及将视觉和触觉输入与动态环境和任务目标相关联的复杂性。为解决这些挑战,我们提出多模态对比无监督强化学习(M2CURL)。该方法采用一种新颖的多模态自监督学习技术,可学习高效的特征表示并促进强化学习算法的更快收敛。我们的方法与具体强化学习算法无关,因此可集成任意现有强化学习算法。我们在Tactile Gym 2仿真器上评估了M2CURL,结果表明该方法在不同操作任务中显著提升了学习效率。与未采用本特征表示学习方法的传统强化学习算法相比,其表现为更快的收敛速度和更高的每回合累积奖励。