A widely used technique for improving policies is success conditioning, in which one collects trajectories, identifies those that achieve a desired outcome, and updates the policy to imitate the actions taken along successful trajectories. This principle appears under many names -- rejection sampling with SFT, goal-conditioned RL, Decision Transformers -- yet what optimization problem it solves, if any, has remained unclear. We prove that success conditioning exactly solves a trust-region optimization problem, maximizing policy improvement subject to a $χ^2$ divergence constraint whose radius is determined automatically by the data. This yields an identity: relative policy improvement, the magnitude of policy change, and a quantity we call action-influence -- measuring how random variation in action choices affects success rates -- are exactly equal at every state. Success conditioning thus emerges as a conservative improvement operator. Exact success conditioning cannot degrade performance or induce dangerous distribution shift, but when it fails, it does so observably, by hardly changing the policy at all. We apply our theory to the common practice of return thresholding, showing this can amplify improvement, but at the cost of potential misalignment with the true objective.


翻译:一种广泛使用的策略改进技术是成功条件化,即收集轨迹,识别那些实现期望结果的轨迹,并更新策略以模仿成功轨迹中的动作。这一原则以多种名称出现——基于SFT的拒绝采样、目标条件化强化学习、决策Transformer——然而,它(如果存在的话)究竟求解了何种优化问题至今仍不清楚。我们证明成功条件化确切地求解了一个信任域优化问题,即在由数据自动确定的χ²散度约束下最大化策略改进。这得出一个恒等式:相对策略改进、策略变化幅度,以及我们称为动作影响(衡量动作选择中的随机变化如何影响成功率的量)的量,在每个状态下均精确相等。因此,成功条件化表现为一个保守的改进算子。精确的成功条件化不会降低性能或引发危险的分布偏移,但当它失效时,其表现是可观测的,即几乎不改变策略。我们将理论应用于常见的回报阈值设定实践,表明这可以放大改进,但代价是可能与真实目标产生偏差。

0
下载
关闭预览

相关内容

《战略战术化:一项综合性述评》
专知会员服务
7+阅读 · 8月5日
《决策中的生成模型:综述》
专知会员服务
49+阅读 · 2025年2月26日
【普林斯顿博士论文】高效决策背后的结构化表征
专知会员服务
40+阅读 · 2024年11月26日
作战决策优势的核心
专知会员服务
95+阅读 · 2023年11月2日
【干货书】计算优化:实践中的成功,415页pdf
专知会员服务
71+阅读 · 2022年12月29日
基于模型的强化学习综述
专知
42+阅读 · 2022年7月13日
强化学习《奖励函数设计: Reward Shaping》详细解读
深度强化学习实验室
20+阅读 · 2020年9月1日
探索(Exploration)还是利用(Exploitation)?强化学习如何tradeoff?
深度强化学习实验室
13+阅读 · 2020年8月23日
强化学习的两大话题之一,仍有极大探索空间
AI科技评论
22+阅读 · 2020年8月22日
国家自然科学基金
43+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
20+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
2+阅读 · 2014年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
Arxiv
0+阅读 · 6月6日
VIP会员
最新内容
博士论文 | 大动作空间中的在线与离线策略学习
论文 | 全球负责任 AI 指数 2026 方法论
专知会员服务
2+阅读 · 8月23日
超致命战场空间中的战术通信生存能力
专知会员服务
2+阅读 · 8月23日
综述 | 面向机器人的视觉触觉智能
专知会员服务
5+阅读 · 8月22日
论文 | 基础模型驱动具身智能体安全综述
专知会员服务
3+阅读 · 8月22日
伊朗-美国-以色列冲突教训
专知会员服务
3+阅读 · 8月22日
《美国陆军 · 战役规划手册(2027年)》237页
专知会员服务
9+阅读 · 8月22日
综述 | 自我中心视频中的视觉语言模型
专知会员服务
3+阅读 · 8月21日
相关基金
国家自然科学基金
43+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
20+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
2+阅读 · 2014年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员