As AI capabilities increasingly surpass human proficiency in complex tasks, current alignment techniques, including SFT and RLHF, face fundamental challenges in ensuring reliable oversight. These methods rely on direct human assessment and become impractical when AI outputs exceed human cognitive thresholds. In response to this challenge, we explore two hypotheses: (1) \textit{Critique of critique can be easier than critique itself}, extending the widely-accepted observation that verification is easier than generation to the critique domain, as critique itself is a specialized form of generation; (2) \textit{This difficulty relationship holds recursively}, suggesting that when direct evaluation is infeasible, performing higher-order critiques (e.g., critique of critique of critique) offers a more tractable supervision pathway. We conduct Human-Human, Human-AI, and AI-AI experiments to investigate the potential of recursive self-critiquing for AI supervision. Our results highlight recursive critique as a promising approach for scalable AI oversight.


翻译:随着人工智能在复杂任务上的能力日益超越人类水平,包括监督微调(SFT)和基于人类反馈的强化学习(RLHF)在内的现有对齐技术,在确保可靠监督方面面临根本性挑战。这些方法依赖于直接的人类评估,当人工智能的输出超出人类认知阈值时,便不再适用。针对这一挑战,我们探讨了两个假设:(1)\textit{对批判的批判可能比批判本身更容易},这一假设将“验证比生成更容易”这一广为接受的观察推广至批判领域,因为批判本身是一种特殊形式的生成;(2)\textit{这种难度关系具有递归性},这意味着当直接评估不可行时,执行更高阶的批判(例如,对批判的批判的批判)提供了一条更可行的监督路径。我们通过人-人、人-人工智能以及人工智能-人工智能实验,研究了递归自批判在人工智能监督中的潜力。我们的结果表明,递归批判是实现可扩展人工智能监督的一种有前景的方法。

0
下载
关闭预览

相关内容

【伯克利博士论文】超越人类监督的视觉智能
专知会员服务
28+阅读 · 2025年8月12日
《重新思考战斗人工智能和人类监督》
专知会员服务
87+阅读 · 2024年5月5日
基于人工反馈的强化学习综述
专知会员服务
66+阅读 · 2023年12月25日
对比自监督学习
深度学习自然语言处理
35+阅读 · 2020年7月15日
【综述】医疗可解释人工智能综述论文
专知
33+阅读 · 2019年7月18日
【强化学习】强化学习+深度学习=人工智能
产业智能官
55+阅读 · 2017年8月11日
国家自然科学基金
3+阅读 · 2017年12月31日
国家自然科学基金
44+阅读 · 2015年12月31日
国家自然科学基金
9+阅读 · 2015年12月31日
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
17+阅读 · 2015年12月31日
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
6+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
国家自然科学基金
12+阅读 · 2013年12月31日
国家自然科学基金
23+阅读 · 2009年12月31日
Arxiv
0+阅读 · 2月10日
VIP会员
最新内容
《人工智能赋能的适应性多功能电磁战》
专知会员服务
9+阅读 · 9月29日
俄乌战场实验室:全面战争如何重塑现代作战
专知会员服务
6+阅读 · 9月29日
2026年美空军协会会议上的无人机系统趋势
专知会员服务
8+阅读 · 9月28日
反制无人机:乌克兰提供的五点启示
专知会员服务
13+阅读 · 9月23日
《各指挥层级均亟需红队能力》报告
专知会员服务
10+阅读 · 9月23日
《航电任务系统框架(FAMOS)》50页报告
专知会员服务
9+阅读 · 9月22日
《对抗行动中的人工智能与自主性》智库报告
专知会员服务
14+阅读 · 9月22日
《从数据到胜利:战争中的分析优势之争》
专知会员服务
17+阅读 · 9月22日
相关VIP内容
【伯克利博士论文】超越人类监督的视觉智能
专知会员服务
28+阅读 · 2025年8月12日
《重新思考战斗人工智能和人类监督》
专知会员服务
87+阅读 · 2024年5月5日
基于人工反馈的强化学习综述
专知会员服务
66+阅读 · 2023年12月25日
相关基金
国家自然科学基金
3+阅读 · 2017年12月31日
国家自然科学基金
44+阅读 · 2015年12月31日
国家自然科学基金
9+阅读 · 2015年12月31日
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
17+阅读 · 2015年12月31日
国家自然科学基金
5+阅读 · 2015年12月31日
国家自然科学基金
6+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
国家自然科学基金
12+阅读 · 2013年12月31日
国家自然科学基金
23+阅读 · 2009年12月31日
Top
微信扫码咨询专知VIP会员