Cancer treatment is at the core a sequential decision-making problem with partial observability, latent patient heterogeneity, and explicit constraints on the budget for medical measurements. Unlike standard Reinforcement Learning (RL) approaches that control state trajectories, cancer treatments permanently modify patients' transition dynamics, changing how states evolve over time. We model cancer treatment as a belief-space planning problem using active inference, deriving an expected free-energy objective that unifies goal-directed control and information acquisition under measurement budgets without. We implement this framework using real clinical cancer data from the AACR Project GENIE Biopharma Collaborative dataset. Results on clinical data demonstrate a simultaneous patient categorization and high treatment efficacy, under real measurement and treatment constraints.
翻译:癌症治疗本质上是部分可观测、患者异质性潜在且医疗测量预算存在显式约束的序贯决策问题。与标准强化学习(RL)方法控制状态轨迹不同,癌症治疗会永久改变患者的转移动力学特性,即改变状态随时间演化的方式。我们采用主动推理将癌症治疗建模为信念空间规划问题,推导出无需测量预算约束即可统一目标导向控制与信息获取的期望自由能目标函数。基于AACR Project GENIE生物制药协作数据集中的真实临床癌症数据实现该框架。临床数据结果表明,在真实测量与治疗约束条件下,该方法能同时实现患者分类与高效治疗。