Reinforcement Learning with Hidden Markov Models for Discovering Decision-Making Dynamics

Major depressive disorder (MDD) presents challenges in diagnosis and treatment due to its complex and heterogeneous nature. Emerging evidence indicates that reward processing abnormalities may serve as a behavioral marker for MDD. To measure reward processing, patients perform computer-based behavioral tasks that involve making choices or responding to stimulants that are associated with different outcomes. Reinforcement learning (RL) models are fitted to extract parameters that measure various aspects of reward processing to characterize how patients make decisions in behavioral tasks. Recent findings suggest the inadequacy of characterizing reward learning solely based on a single RL model; instead, there may be a switching of decision-making processes between multiple strategies. An important scientific question is how the dynamics of learning strategies in decision-making affect the reward learning ability of individuals with MDD. Motivated by the probabilistic reward task (PRT) within the EMBARC study, we propose a novel RL-HMM framework for analyzing reward-based decision-making. Our model accommodates learning strategy switching between two distinct approaches under a hidden Markov model (HMM): subjects making decisions based on the RL model or opting for random choices. We account for continuous RL state space and allow time-varying transition probabilities in the HMM. We introduce a computationally efficient EM algorithm for parameter estimation and employ a nonparametric bootstrap for inference. We apply our approach to the EMBARC study to show that MDD patients are less engaged in RL compared to the healthy controls, and engagement is associated with brain activities in the negative affect circuitry during an emotional conflict task.

翻译：重度抑郁症（MDD）因其复杂异质性在诊断和治疗方面面临挑战。新近证据表明，奖赏处理异常可能作为MDD的行为标志。为测量奖赏处理，患者需完成基于计算机的行为任务，包括对关联不同结果的刺激做出选择或反应。通过拟合强化学习（RL）模型提取参数，可量化奖赏处理的多个维度，进而表征患者在行为任务中的决策模式。最新研究发现，仅基于单一RL模型刻画奖赏学习存在局限性——决策过程可能在多种策略间切换。重要科学问题在于：决策过程中学习策略的动态变化如何影响MDD患者的奖赏学习能力？受EMBARC研究中概率性奖赏任务（PRT）启发，我们提出新颖的RL-HMM框架分析基于奖赏的决策。该模型在隐马尔可夫模型（HMM）框架下容纳两种学习策略的切换：被试依据RL模型决策或随机选择。我们考虑了连续RL状态空间，允许HMM中时变转移概率，并引入计算高效的EM算法进行参数估计，采用非参数自助法进行推断。将本方法应用于EMBARC研究表明：与健康对照组相比，MDD患者的RL参与度降低，且该参与度与情绪冲突任务中负性情感环路的大脑活动相关。

相关内容

隐马尔科夫模型

关注 18

隐马儿可夫模型：HMM，hidden Markov model，是可用于标注问题的统计学习模型，描述由隐藏的马尔可夫链随机生成观测序列的过程，属于生成模型。隐马尔可夫模型是关于时序的概率模型，描述由一个隐藏的马尔可夫链随机生成不可观测的状态随机序列，再有各个状态生成一个观测而产生观测随机序列的过程。隐藏的马尔可夫链随机生成的状态的序列，称为状态序列。每个状态生成一个观测，而由此产生的观测的随机序列，称为观测序列。

《生成式模型: 变分自编码器与扩散模型》，75页ppt，Google DeepMind科学家Ruiqi Gao

专知会员服务

66+阅读 · 2023年6月10日

UCM《机器学习导论笔记》，80页pdf CSE176 Introduction to Machine Learning

专知会员服务

32+阅读 · 2021年9月29日

FlowQA: Grasping Flow in History for Conversational Machine Comprehension

专知会员服务

34+阅读 · 2019年10月18日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

50+阅读 · 2019年10月17日