How is it that humans can solve complex planning tasks so efficiently despite limited cognitive resources? One reason is its ability to know how to use its limited computational resources to make clever choices. We postulate that people learn this ability from trial and error (metacognitive reinforcement learning). Here, we systematize models of the underlying learning mechanisms and enhance them with more sophisticated additional mechanisms. We fit the resulting 86 models to human data collected in previous experiments where different phenomena of metacognitive learning were demonstrated and performed Bayesian model selection. Our results suggest that a gradient ascent through the space of cognitive strategies can explain most of the observed qualitative phenomena, and is therefore a promising candidate for explaining the mechanism underlying metacognitive learning.
翻译:人类如何在认知资源有限的情况下高效解决复杂的规划任务?其中一个原因是人类具备知道如何利用有限计算资源做出明智选择的能力。我们假设这种能力是通过试错习得的(元认知强化学习)。在此,我们系统化了潜在学习机制的模型,并通过更复杂的附加机制对其进行增强。我们将最终得到的86个模型与先前实验中收集的人类数据进行拟合——这些实验展示了元认知学习的不同现象——并进行了贝叶斯模型选择。结果表明,通过认知策略空间的梯度上升可以解释大部分观察到的定性现象,因此是解释元认知学习背后机制的一个有前景的候选方案。