Markdown skill libraries for LLM agents ship as free-form prose, forcing the agent to re-derive both the input schema and the concrete invocation syntax on every retrieval. We observe that this often produces a "confused -> re-retrieve -> still confused" loop in which the agent issues a partially-correct action, receives uninformative environment feedback, and re-retrieves the same prose. We propose Skill-as-Pseudocode (SaP), an automatic conversion of markdown skill libraries into typed pseudocode with deterministic quality control. For each cluster of similar procedural passages drawn from one or more skills, SaP extracts a typed contract and filters it through a four-check deterministic verifier (coverage, binding, replacement, risk). Promoted contracts are inlined into a rewritten skill skeleton together with restored concrete action templates, giving the agent two complementary signals: a typed signature for what the skill does and a concrete template for how to invoke it. On the 134-game ALFWorld unseen split with gpt-4o-mini, pooled across three seeds, SaP wins 82/402 paired games versus 47/402 for the Graph-of-Skills (GoS) baseline (pooled McNemar p = 8.2e-5), at -22.8 +/- 6.4% input tokens and -14.5 +/- 4.1% LLM calls per game.


翻译:面向LLM智能体的Markdown技能库以自由格式文本形式提供,迫使智能体在每次检索时重新推导输入模式与具体调用语法。我们观察到这常导致"困惑→重新检索→依旧困惑"的循环:智能体发出部分正确的动作后收到无信息量的环境反馈,继而重新检索同一段文本。为此提出技能即伪代码(Skill-as-Pseudocode, SaP)方法,将Markdown技能库自动转换为带有确定性质量控制的类型化伪代码。针对源自一个或多个技能的相似过程性文本聚类,SaP提取类型化契约并通过四重确定性检验器(覆盖性、绑定性、替代性、风险性)进行过滤。被通过的契约将内联至重写的技能骨架中,同时保留具体动作模板,从而为智能体提供两种互补信号:描述技能功能的类型化签名,以及指示调用方法的具体模板。在包含134个游戏的ALFWorld未见分片测试集上(使用GPT-4o-mini,三组随机种子合并统计),SaP在402场配对游戏中获胜82场,而Graph-of-Skills(GoS)基线仅获胜47场(合并McNemar检验p=8.2e-5),同时每场游戏输入Token减少22.8±6.4%,LLM调用次数减少14.5±4.1%。

0
下载
关闭预览

相关内容

智能体技能综合综述:分类、技术与应用
专知会员服务
35+阅读 · 5月11日
LLM/智能体作为数据分析师:综述
专知会员服务
38+阅读 · 2025年9月30日
ML、DL、NLP面试常考知识点、代码、算法理论基础汇总分享
【NLP】万字长文概述NLP中的深度学习技术
产业智能官
18+阅读 · 2019年7月7日
面试题:文本摘要中的NLP技术
七月在线实验室
15+阅读 · 2019年5月13日
NLP-Progress记录NLP最新数据集、论文和代码: 助你紧跟NLP前沿
中国人工智能学会
12+阅读 · 2018年11月15日
深度学习文本分类方法综述(代码)
中国人工智能学会
28+阅读 · 2018年6月16日
国家自然科学基金
4+阅读 · 2017年12月31日
国家自然科学基金
18+阅读 · 2017年12月31日
国家自然科学基金
6+阅读 · 2017年12月31日
国家自然科学基金
13+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
5+阅读 · 2014年12月31日
VIP会员
最新内容
《无人机脆弱性利用:网络空间力量的新域》
专知会员服务
2+阅读 · 8月1日
美空军如何将人工智能从战场部署至后方机关
专知会员服务
11+阅读 · 7月31日
《史诗怒火行动:多域前瞻评估》49页报告
专知会员服务
7+阅读 · 7月31日
《英国防部:未来空战系统数字化战略》33页
专知会员服务
5+阅读 · 7月31日
《面向自主飞行网络的智能体人工智能架构》
专知会员服务
7+阅读 · 7月31日
“史诗怒火”行动:现代多域作战的重要节点
专知会员服务
8+阅读 · 7月30日
《下一代无线网络中的多无人机通信资源管理》
相关资讯
相关基金
国家自然科学基金
4+阅读 · 2017年12月31日
国家自然科学基金
18+阅读 · 2017年12月31日
国家自然科学基金
6+阅读 · 2017年12月31日
国家自然科学基金
13+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
5+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员