Deep reinforcement learning (RL) algorithms enable the development of fully autonomous agents that can interact with the environment. Brain-computer interface (BCI) systems decipher human implicit brain signals regardless of the explicit environment. In this study, we integrated deep RL and BCI to improve beneficial human interventions in autonomous systems and the performance in decoding brain activities by considering environmental factors. Shared autonomy was allowed between the action command decoded from the electroencephalography (EEG) of the human agent and the action generated from the twin delayed DDPG (TD3) agent for a given environment. Our proposed copilot control scheme with a full blocker (Co-FB) significantly outperformed the individual EEG (EEG-NB) or TD3 control. The Co-FB model achieved a higher target approaching score, lower failure rate, and lower human workload than the EEG-NB model. The Co-FB control scheme had a higher invisible target score and level of allowed human intervention than the TD3 model. We also proposed a disparity d-index to evaluate the effect of contradicting agent decisions on the control accuracy and authority of the copilot model. We found a significant correlation between the control authority of the TD3 agent and the performance improvement of human EEG classification with respect to the d-index. We also observed that shifting control authority to the TD3 agent improved performance when BCI decoding was not optimal. These findings indicate that the copilot system can effectively handle complex environments and that BCI performance can be improved by considering environmental factors. Future work should employ continuous action space and different multi-agent approaches to evaluate copilot performance.
翻译:摘要:深度强化学习算法能够开发出与环境交互的完全自主智能体。脑机接口系统可破译人类隐性脑信号,无需依赖显式环境。本研究将深度强化学习与脑机接口相结合,通过考虑环境因素,提升自主系统中人类有益干预的有效性及脑活动解码性能。针对给定环境,人类智能体通过脑电图解码的动作指令与双延迟深度确定性策略梯度(TD3)智能体生成的动作之间实现了共享自主。我们提出的带完全阻断器的副驾驶控制方案(Co-FB)在性能上显著优于单独使用脑电图(EEG-NB)或TD3控制。与EEG-NB模型相比,Co-FB模型获得了更高的目标接近得分、更低的失败率和更少的人类工作量。与TD3模型相比,Co-FB控制方案具有更高的隐形目标得分和允许的人类干预水平。我们还提出了差异指数(d-index)来评估智能体矛盾决策对副驾驶模型控制精度与权限的影响。研究发现,TD3智能体的控制权限与人类脑电图分类性能提升之间存在显著相关性(基于d-index)。同时观察到,当脑机接口解码非最优时,将控制权限转移至TD3智能体可提升性能。这些结果表明,副驾驶系统能有效处理复杂环境,且通过考虑环境因素可改善脑机接口性能。未来工作应采用连续动作空间及不同多智能体方法来评估副驾驶性能。