Effective exploration in reinforcement learning requires not only tracking where an agent has been, but also understanding how the agent perceives and represents the world. To learn powerful representations, an agent should actively explore states that contribute to its knowledge of the environment. Temporal representations can capture the information necessary to solve a wide range of potential tasks while avoiding the computational cost associated with full state reconstruction. In this paper, we propose an exploration method that leverages temporal contrastive representations to guide exploration, prioritizing states with unpredictable future outcomes. We demonstrate that such representations can enable the learning of complex exploratory x in locomotion, manipulation, and embodied-AI tasks, revealing capabilities and behaviors that traditionally require extrinsic rewards. Unlike approaches that rely on explicit distance learning or episodic memory mechanisms (e.g., quasimetric-based methods), our method builds directly on temporal similarities, yielding a simpler yet effective strategy for exploration.
翻译:在强化学习中,高效探索不仅需要追踪智能体到过何处,还需理解其如何感知与表征世界。为学习强大表征,智能体应主动探索能增进环境认知的状态。时间表征既能捕捉解决广泛潜在任务所需的信息,又可避免完整状态重建带来的计算开销。本文提出一种利用时间对比表征指导探索的方法,优先关注具有不可预测未来结果的状态。实验证明,该表征可帮助在运动控制、操控任务及具身AI任务中学习复杂的探索行为,展现出传统上需依赖外部奖励才能实现的能力与行为模式。不同于依赖显式距离学习或情景记忆机制(如拟度量方法)的现有方案,本方法直接基于时间相似性构建,形成更简洁高效的探索策略。