Agentic search enables LLMs to solve complex multi-hop questions through iterative reasoning and external search. Despite the effectiveness, these systems often suffer from a critical limitation in practice: agents fail to recognize their own knowledge boundaries, blindly triggering searches when internal knowledge suffices and failing to terminate search even when adequate evidence has been collected. The lack of self-awareness leads to severe \textbf{over-search}, incurring substantial inference latency and prohibitive computational cost. To this end, we propose SAAS, a novel RL framework designed to cultivate dynamic self-awareness that precisely regulates search behavior without compromising accuracy. SAAS introduces three key components: (i) a search boundary modeling mechanism, which identifies the search boundary under the evolving policy by contrasting search-disabled and search-enabled rollouts; (ii) a boundary-aware reward module, which translates this boundary awareness into trajectory-level penalties, suppressing unnecessary and redundant searches; and (iii) a stage-wise optimization strategy, which leverages a sequential curriculum to prioritize reasoning over search regularization, thereby avoiding reward hacking. Extensive experiments demonstrate that SAAS substantially reduces over-search, while maintaining accuracy. Our code and implementation details are released at https://github.com/XMUDeepLIT/SAAS.


翻译:智能体搜索使大语言模型能够通过迭代推理和外部搜索解决复杂的多跳问题。尽管效果显著,但这类系统在实践中常面临关键局限:智能体无法识别自身知识边界,在内部知识足够时盲目触发搜索,即便已收集充分证据仍未能终止搜索过程。这种自感知缺失导致严重的\textbf{过度搜索}现象,引发大量推理延迟与高昂计算成本。为此,我们提出SAAS——一种新型强化学习框架,旨在培养动态自感知能力以精确调控搜索行为且不损害准确性。SAAS包含三大核心组件:(1)搜索边界建模机制:通过对比禁用搜索与启用搜索的轨迹序列,在策略演化过程中识别搜索边界;(2)边界感知奖励模块:将边界感知转化为轨迹级惩罚,抑制不必要及冗余搜索;(3)分阶段优化策略:采用序列化课程学习优先强化推理能力再优化搜索正则化,从而规避奖励破解问题。大量实验表明,SAAS在保持准确性的同时显著减少了过度搜索。我们的代码及实现细节已开源至 https://github.com/XMUDeepLIT/SAAS。

0
下载
关闭预览

相关内容

互联网
KARL:基于强化学习的知识智能体
专知会员服务
14+阅读 · 3月7日
AI 智能体系统:体系架构、应用场景及评估范式
数据驱动的具身学习探索
专知会员服务
11+阅读 · 2025年2月26日
可解释强化学习,Explainable Reinforcement Learning: A Survey
专知会员服务
133+阅读 · 2020年5月14日
深度学习搜索,Exploring Deep Learning for Search
专知会员服务
62+阅读 · 2020年5月9日
PlaNet 简介:用于强化学习的深度规划网络
谷歌开发者
13+阅读 · 2019年3月16日
展望:模型驱动的深度学习
人工智能学家
12+阅读 · 2018年1月23日
【强化学习】强化学习+深度学习=人工智能
产业智能官
55+阅读 · 2017年8月11日
国家自然科学基金
3+阅读 · 2017年12月31日
国家自然科学基金
1+阅读 · 2016年12月31日
国家自然科学基金
43+阅读 · 2015年12月31日
国家自然科学基金
52+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
13+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
18+阅读 · 2009年12月31日
国家自然科学基金
17+阅读 · 2008年12月31日
VIP会员
最新内容
俄乌无人机战争的六大启示
专知会员服务
9+阅读 · 8月3日
《无人机空中监控:通信实验洞察》
专知会员服务
6+阅读 · 8月3日
从采集到决策:美军视角下的战术情报范式重构
《履带式无人地面战车技术发展现状》
专知会员服务
6+阅读 · 8月2日
《无人机脆弱性利用:网络空间力量的新域》
专知会员服务
8+阅读 · 8月1日
美空军如何将人工智能从战场部署至后方机关
专知会员服务
14+阅读 · 7月31日
相关VIP内容
相关基金
国家自然科学基金
3+阅读 · 2017年12月31日
国家自然科学基金
1+阅读 · 2016年12月31日
国家自然科学基金
43+阅读 · 2015年12月31日
国家自然科学基金
52+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
13+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
18+阅读 · 2009年12月31日
国家自然科学基金
17+阅读 · 2008年12月31日
Top
微信扫码咨询专知VIP会员