We study LLM policy synthesis: using a language model to iteratively generate programmatic agent policies for multi-agent environments. Rather than training neural policies via reinforcement learning, our framework prompts an LLM to produce Python policy functions, evaluates them in self-play, and refines them using performance feedback across iterations. We investigate feedback engineering (the design of what evaluation information is shown to the LLM during refinement) comparing sparse feedback (scalar reward only) against dense feedback (reward plus social metrics: efficiency, equality, sustainability, peace). Across two canonical Sequential Social Dilemmas (Gathering and Cleanup) and two frontier LLMs (Claude Sonnet 4.6, Gemini 3.1 Pro), dense feedback consistently matches or exceeds sparse feedback on all metrics. We explain the asymmetry through feedback aliasing: when scalar reward alone maps distinct failure modes to the same value (e.g., under- vs. over-cleaning), social metrics break the alias and let the LLM diagnose which corrective direction to take. Social metrics thus function as a coordination signal rather than a distraction, yielding strategies such as Voronoi territory partitioning and waste-adaptive cleaner schedules. Code at https://github.com/vicgalle/llm-policies-social-dilemmas.


翻译:我们研究大语言模型策略生成:利用语言模型为多智能体环境迭代式地生成程序化智能体策略。我们的框架不采用强化学习来训练神经策略,而是通过提示大语言模型生成Python策略函数,在自博弈中评估这些函数,并依据跨迭代的性能反馈进行优化。我们研究了反馈工程(即设计在优化过程中向大语言模型展示何种评估信息),比较了稀疏反馈(仅基于标量奖励)与密集反馈(奖励加社会指标:效率、平等、可持续性、和平)。在两个经典序贯社会困境(Gathering和Cleanup)及两个前沿大语言模型(Claude Sonnet 4.6、Gemini 3.1 Pro)上,密集反馈在所有指标上一致达到或超过稀疏反馈的性能。我们通过反馈混叠解释这一不对称现象:当标量奖励将不同失败模式映射至相同数值时(例如清理不足与清理过度),社会指标能打破混叠,使大语言模型能够诊断应采取的校正方向。因此,社会指标起到协调信号而非干扰作用,产生了Voronoi区域划分和废弃物自适应清理调度等策略。代码地址:https://github.com/vicgalle/llm-policies-social-dilemmas。

0
下载
关闭预览

相关内容

《多智能体大语言模型系统的可靠决策研究》
专知会员服务
41+阅读 · 2月2日
深度强化学习中的奖励模型:综述
专知会员服务
29+阅读 · 2025年6月20日
大语言模型在规划与调度问题上的应用
专知会员服务
54+阅读 · 2025年1月12日
【伯克利博士论文】以人为中心的奖励设计
专知会员服务
28+阅读 · 2024年9月23日
大语言模型视角下的智能规划方法综述
专知会员服务
139+阅读 · 2024年4月20日
强化学习《奖励函数设计: Reward Shaping》详细解读
深度强化学习实验室
20+阅读 · 2020年9月1日
PlaNet 简介:用于强化学习的深度规划网络
谷歌开发者
13+阅读 · 2019年3月16日
超全总结:神经网络加速之量化模型 | 附带代码
国家自然科学基金
43+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
6+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
VIP会员
最新内容
驱动军事决策变革的顶尖人工智能指挥系统
专知会员服务
7+阅读 · 8月11日
非对称防御中的自组织临界性:俄乌战争
专知会员服务
10+阅读 · 8月10日
《战争中的大语言模型监管》
专知会员服务
13+阅读 · 8月10日
《边缘计算关键技术分析及美军作战实践应用》
边缘计算的军事应用
专知会员服务
12+阅读 · 8月9日
一种考虑资源机动性的武器目标分配混合算法
专知会员服务
13+阅读 · 8月8日
相关VIP内容
相关基金
国家自然科学基金
43+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
6+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员