Humans must flexibly arbitrate between exploring alternatives and exploiting learned strategies, yet they frequently exhibit maladaptive persistence by continuing to execute failing strategies despite accumulating negative evidence. Here we propose a ``confidence-freeze'' account that reframes such persistence as a dynamic learning state rather than a stable dispositional trait. Using a multi-reversal two-armed bandit task across three experiments (total N = 332; 19,920 trials), we first show that human learners normally make use of the symmetric statistical structure inherent in outcome trajectories: runs of successes provide positive evidence for environmental stability and thus for strategy maintenance, whereas runs of failures provide negative evidence and should raise switching probability. Behaviour in the control group conformed to this normative pattern. However, individuals who experienced a high rate of early success (90\% vs.\ 60\%) displayed a robust and selective distortion after the first reversal: they persisted through long stretches of non-reward (mean = 6.2 consecutive losses) while their metacognitive confidence ratings simultaneously dropped from 5 to 2 on a 7-point scale.
翻译:人类必须灵活地在探索替代方案与利用已学策略之间进行仲裁,然而他们常常表现出适应不良的持续性行为——即使不断积累负面证据,仍继续执行失败的策略。本文提出“自信冻结”解释,将此类持续性行为重新定义为动态学习状态而非稳定的特质倾向。通过三项实验(总样本量N=332;共19,920次试验)中采用多重反转双臂老虎机任务,我们首先证明人类学习者通常会利用结果轨迹中固有的对称统计结构:成功序列为环境稳定性提供正面证据,从而支持策略维持;而失败序列则提供负面证据,理应提高切换概率。对照组的行为符合这一规范模式。然而,经历早期高成功率(90%对比60%)的个体在首次反转后表现出显著的选择性失真:他们持续经历长串无奖励时段(平均连续6.2次损失),同时其元认知自信评分从7分量表的5分骤降至2分。