Agentic reinforcement learning (RL) enables LLM agents to improve continuously from environment rewards, yet the resulting policies do not systematically accumulate reusable strategies that generalize across tasks. Modular skills can provide such reusable strategies, yet existing skill-augmented RL methods decouple skill creation from policy optimization, risking adopting skills that conflict with the evolving policy. Inspired by Anthropic's Skill Creator, we introduce ReSkill, an RL-in-the-loop skill creation framework that reconciles skill evolution with policy learning. ReSkill exploits the group-wise structure of GRPO to naturally embed three mechanisms with only marginal additional overhead: (1) an assertion-driven skill creator that diagnoses failures from past experience and proposes conditional, trigger-based skill revisions; (2) within-group rollout sampling that enables controlled comparison of skill versions, capturing which version best supports the policy's ongoing learning; and (3) Thompson Sampling with adaptive discounting to balance exploration and exploitation in skill version selection as the policy evolves. Across several domains, ReSkill consistently outperforms existing memory and skill-based RL methods, with the largest gains on unseen tasks. Analysis of the skill lifecycle shows skills being automatically created, tested, refined, and pruned as the policy improves, demonstrating reconciled skill-policy co-evolution.


翻译:智能体强化学习使大语言模型智能体能够通过环境奖励持续改进,然而由此产生的策略无法系统积累可跨任务泛化的可复用策略。模块化技能可提供此类可复用策略,但现有技能增强型强化学习方法将技能创建与策略优化相分离,存在采纳与演进策略冲突技能的风险。受Anthropic的Skill Creator启发,我们提出ReSkill——一种将技能演化与策略学习相结合的强化学习闭环技能创建框架。ReSkill利用GRPO的分组结构特性,仅需微量额外开销即可自然嵌入三种机制:(1)基于断言的技能创建器,通过诊断过往经验中的失败模式提出条件触发式技能修正方案;(2)组内轨迹采样实现技能版本的受控对比,精准识别最适配策略持续学习的技能版本;(3)结合自适应折扣的汤普森采样,在策略演进过程中平衡技能版本选择的探索与利用。在多个领域中,ReSkill始终优于现有基于记忆和技能的强化学习方法,尤其在未见任务上表现突出。对技能生命周期的分析表明,随着策略优化,技能能够自动完成创建、测试、优化与裁剪,验证了技能与策略协同演化的可行性。

0
下载
关闭预览

相关内容

Agentic RL:框架、实践与长程智能体训练
专知会员服务
23+阅读 · 6月24日
《基于分层多智能体强化学习的逼真空战协同策略》
专知会员服务
48+阅读 · 2025年10月30日
【NTU博士论文】基于协作式多智能体强化学习的决策制定
面向关系建模的合作多智能体深度强化学习综述
专知会员服务
42+阅读 · 2025年4月18日
基于学习机制的多智能体强化学习综述
专知会员服务
64+阅读 · 2024年4月16日
基于多智能体强化学习的协同目标分配
专知会员服务
142+阅读 · 2023年9月5日
「基于通信的多智能体强化学习」 进展综述
【综述】多智能体强化学习算法理论研究
深度强化学习实验室
17+阅读 · 2020年9月9日
强化学习《奖励函数设计: Reward Shaping》详细解读
深度强化学习实验室
20+阅读 · 2020年9月1日
探索(Exploration)还是利用(Exploitation)?强化学习如何tradeoff?
深度强化学习实验室
13+阅读 · 2020年8月23日
强化学习的两大话题之一,仍有极大探索空间
AI科技评论
22+阅读 · 2020年8月22日
【强化学习】强化学习+深度学习=人工智能
产业智能官
55+阅读 · 2017年8月11日
国家自然科学基金
43+阅读 · 2015年12月31日
国家自然科学基金
41+阅读 · 2015年12月31日
国家自然科学基金
24+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
10+阅读 · 2013年12月31日
国家自然科学基金
18+阅读 · 2009年12月31日
国家自然科学基金
23+阅读 · 2009年12月31日
国家自然科学基金
50+阅读 · 2009年12月31日
国家自然科学基金
17+阅读 · 2008年12月31日
国家自然科学基金
12+阅读 · 2008年12月31日
VIP会员
最新内容
边缘计算的军事应用
专知会员服务
6+阅读 · 8月9日
一种考虑资源机动性的武器目标分配混合算法
专知会员服务
8+阅读 · 8月8日
《多域冲突比较支持模型》60页
专知会员服务
13+阅读 · 8月7日
相关基金
国家自然科学基金
43+阅读 · 2015年12月31日
国家自然科学基金
41+阅读 · 2015年12月31日
国家自然科学基金
24+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
10+阅读 · 2013年12月31日
国家自然科学基金
18+阅读 · 2009年12月31日
国家自然科学基金
23+阅读 · 2009年12月31日
国家自然科学基金
50+阅读 · 2009年12月31日
国家自然科学基金
17+阅读 · 2008年12月31日
国家自然科学基金
12+阅读 · 2008年12月31日
Top
微信扫码咨询专知VIP会员