In recent years, generative AI (GenAI) in educational settings has become ubiquitous in university students' daily lives, despite its potential to induce over-reliance, metacognitive disengagement, and diminished learning when used unrestrictedly. While most prior research has focused on how to pedagogically scaffold its usage, the question of when to allow off-the-shelf GenAI remains understudied and lacks pedagogically grounded empirical investigation. We treat access timing itself as a form of implicit scaffolding and operationalize it through a reinforcement learning (RL) agent that decides when students should access GenAI, with a reward function grounded in metacognitive theory, cognitive load theory, and productive failure. In a mixed-methods controlled lab study with N=105 higher education students, we compared the agent's effect on learning gains and metacognitive engagement to unrestricted and fully restricted use. Results show that strategically timed GenAI access under the reinforcement learning condition improved objective post-test performance and metacognitive accuracy compared with unrestricted access, while reducing task errors and time on task relative to complete withholding, thus outperforming both approaches without the need for explicit metacognitive prompts or structured scaffolding. However, no between-condition differences emerged on self-reported metacognitive awareness. Overall, timing of GenAI access therefore is a tractable, theoretically grounded, and scalable pedagogical strategy that improves over completely unrestricted and withheld access, compatible with off-the-shelf tools and potentially low adoption barrier. This opens up a new research area that explores how access timing can be facilitated by educators and implemented in human-AI learning system design.
翻译:近年来,生成式AI(GenAI)在高等教育场景中已成为大学生日常生活的普遍存在,尽管在无限制使用下,它可能引发过度依赖、元认知脱离以及学习效果下降等问题。尽管以往研究主要聚焦于如何从教学法上为其使用提供脚手架支持,关于何时允许学生接入现成的GenAI工具这一议题,仍缺乏基于教学理论的实证探究。我们将访问时机本身视为一种隐含的脚手架,并通过强化学习(RL)智能体将其操作化——该智能体决定学生何时应当使用GenAI,其奖励函数基于元认知理论、认知负荷理论与生产性失败理论构建。在一项包含105名高等教育学生的混合方法受控实验室研究中,我们比较了该智能体对学习收益及元认知参与的影响,与无限制使用和完全限制使用两种条件进行对比。结果表明,相较于无限制访问,强化学习条件下策略性安排的GenAI访问时机提升了客观后测成绩及元认知准确性;同时相较于完全禁用,减少了任务错误与任务时间,从而在无需显式元认知提示或结构化脚手架的情况下,超越了这两种方法。然而,在自我报告的元认知意识方面,各组间未发现显著差异。总体而言,GenAI访问时机是一种可操作、有理论依据且可扩展的教学策略,优于完全开放或完全禁止的访问方式,且与现成工具兼容并具有较低的采纳门槛。这开辟了一个新的研究领域,探索教育者如何促进访问时机的设计,以及如何在人机协同学习系统中实现这一策略。