Active Inference is a recent framework for modeling planning under uncertainty. Empirical and theoretical work have now begun to evaluate the strengths and weaknesses of this approach and how it might be improved. A recent extension - the sophisticated inference (SI) algorithm - improves performance on multi-step planning problems through recursive decision tree search. However, little work to date has been done to compare SI to other established planning algorithms. SI was also developed with a focus on inference as opposed to learning. The present paper has two aims. First, we compare performance of SI to Bayesian reinforcement learning (RL) schemes designed to solve similar problems. Second, we present an extension of SI - sophisticated learning (SL) - that more fully incorporates active learning during planning. SL maintains beliefs about how model parameters would change under the future observations expected under each policy. This allows a form of counterfactual retrospective inference in which the agent considers what could be learned from current or past observations given different future observations. To accomplish these aims, we make use of a novel, biologically inspired environment designed to highlight the problem structure for which SL offers a unique solution. Here, an agent must continually search for available (but changing) resources in the presence of competing affordances for information gain. Our simulations show that SL outperforms all other algorithms in this context - most notably, Bayes-adaptive RL and upper confidence bound algorithms, which aim to solve multi-step planning problems using similar principles (i.e., directed exploration and counterfactual reasoning). These results provide added support for the utility of Active Inference in solving this class of biologically-relevant problems and offer added tools for testing hypotheses about human cognition.
翻译:主动推理是近年来用于建模不确定性下规划的新框架。实证与理论工作已开始评估该方法的优劣及其改进方向。近期扩展——深度推理(SI)算法——通过递归决策树搜索提升了多步规划问题的性能。然而,目前鲜有研究将SI与其他成熟规划算法进行对比。此外,SI的设计侧重推理而非学习。本文有两项目标:第一,将SI的性能与旨在解决类似问题的贝叶斯强化学习(RL)方案进行对比;第二,提出SI的扩展——深度学习(SL)——更全面地整合规划过程中的主动学习。SL会维持关于模型参数在不同策略预期未来观测下如何变化的信念,这使得一种反事实回溯推理成为可能——智能体可根据不同未来观测,考量从当前或历史观测中能学到什么。为实现这些目标,我们采用了一种受生物学启发的新型环境,其突出体现SL能提供独特解决方案的问题结构:智能体需在存在竞争性信息获取线索的情况下,持续搜索可获取(但动态变化的)资源。模拟结果表明,在该场景中SL的表现优于所有其他算法——尤其优于贝叶斯自适应RL和置信上界算法(后者旨在通过相似原理(即定向探索与反事实推理)解决多步规划问题)。这些结果进一步支持主动推理在解决此类生物学相关问题时的实用性,并为检验人类认知假说提供了新工具。