Reinforcement learning has achieved remarkable success in learning complex control policies, yet its applicability remains limited due to sample inefficiency and poor generalization across tasks. In this work, we propose RepMT-SAC, a framework for multi-task RL that enables efficient knowledge sharing and robust transfer to new tasks. RepMT-SAC uses spectral MDP decomposition to capture transferable dynamics, structuring the value function into a task-agnostic core with a minimal task-specific adjustment. This design allows for strong zero-shot performance on in-distribution tasks and rapid few-shot adaptation to out-of-distribution tasks. We evaluate RepMT-SAC on quadcopter trajectory-following tasks across in-distribution and out-of-distribution contexts, demonstrating that it outperforms baselines by up to 30%.
翻译:强化学习在学习复杂控制策略方面取得了显著成功,但由于样本效率低下和跨任务泛化能力不足,其应用仍受到限制。在这项工作中,我们提出了RepMT-SAC——一个多任务强化学习框架,能够实现高效的知识共享和对新任务的鲁棒迁移。RepMT-SAC使用谱MDP分解来捕捉可迁移的动态特性,将价值函数结构化为任务无关的核心部分和最小的任务特定调整。这种设计使得在分布内任务上具备强大的零样本性能,并在分布外任务上实现快速的小样本适应。我们在四旋翼飞行器轨迹跟踪任务上对RepMT-SAC进行了评估,涵盖了分布内和分布外场景,结果表明其性能优于基线方法高达30%。