Driven by the algorithmic advancements in reinforcement learning and the increasing number of implementations of human-AI collaboration, Collaborative Reinforcement Learning (CRL) has been receiving growing attention. Despite this recent upsurge, this area is still rarely systematically studied. In this paper, we provide an extensive survey, investigating CRL methods based on both interactive reinforcement learning algorithms and human-AI collaborative frameworks that were proposed in the past decade. We elucidate and discuss via synergistic analysis methods both the growth of the field and the state-of-the-art; we conceptualise the existing frameworks from the perspectives of design patterns, collaborative levels, parties and capabilities, and review interactive methods and algorithmic models. Specifically, we create a new Human-AI CRL Design Trajectory Map, as a systematic modelling tool for the selection of existing CRL frameworks, as well as a method of designing new CRL systems, and finally of improving future CRL designs. Furthermore, we elaborate generic Human-AI CRL challenges, providing the research community with a guide towards novel research directions. The aim of this paper is to empower researchers with a systematic framework for the design of efficient and 'natural' human-AI collaborative methods, making it possible to work on maximised realisation of humans' and AI's potentials.
翻译:受强化学习算法进展与人机协作应用日益增多的推动,协同强化学习(CRL)正获得越来越多的关注。尽管近年来该领域发展迅速,但系统的研究仍较为匮乏。本文对过去十年间提出的基于交互式强化学习算法与人机协作框架的CRL方法进行了全面综述。我们通过协同分析方法,阐述并探讨了该领域的发展态势及前沿成果;从设计模式、协作层级、参与方与能力等视角对现有框架进行了概念化梳理,并综述了交互方法与算法模型。具体而言,我们构建了一个全新的人机CRL设计轨迹图谱,既可作为现有CRL框架选择的系统建模工具,也可作为设计新型CRL系统的方法论,最终服务于未来CRL设计的改进。此外,我们深入阐述了人机CRL面临的通用挑战,为该领域研究者指引了新颖的研究方向。本文旨在为研究者提供一套系统化框架,用于设计高效且"自然"的人机协同方法,从而充分释放人类与人工智能的潜力。