Cancer treatment is at the core a sequential decision-making problem with partial observability, latent patient heterogeneity, and explicit constraints on the budget for medical measurements. Unlike standard Reinforcement Learning (RL) approaches that control state trajectories, cancer treatments permanently modify patients' transition dynamics, changing how states evolve over time. We model cancer treatment as a belief-space planning problem using active inference, deriving an expected free-energy objective that unifies goal-directed control and information acquisition under measurement budgets without. We implement this framework using real clinical cancer data from the AACR Project GENIE Biopharma Collaborative dataset. Results on clinical data demonstrate a simultaneous patient categorization and high treatment efficacy, under real measurement and treatment constraints.
翻译:癌症治疗本质上是一个具有部分可观测性、潜在患者异质性以及医疗检测预算明确约束的序贯决策问题。与标准强化学习方法控制状态轨迹不同,癌症治疗会永久性地改变患者的转移动力学,从而影响状态随时间演化的方式。我们将癌症治疗建模为基于主动推理的信念空间规划问题,推导出预期自由能目标函数,该函数在预算约束下统一了目标导向控制与信息获取。我们采用AACR项目GENIE生物制药合作数据集的真实临床癌症数据实现该框架。基于临床数据的实验结果表明,在真实检测与治疗约束条件下,该方法能够实现患者分类与高治疗效用的同步优化。