In hierarchical reasoning, failures often originate at intermediate decision points where the agent commits to a wrong branch without recognizing that it lacks critical information. Rather than treating clarification as an external uncertainty trigger, we propose ACTION-RATING, a formulation that places it inside the agent's action space on a shared ordinal scale with navigation, so that asking competes directly with acting at every decision point and help-seeking becomes observable at intermediate states. Two structurally distinct information-seeking modes emerge from the agent's own ratings: mandatory (no viable branch) and opportunistic (residual uncertainty despite a leading candidate). On Harmonized Tariff Schedule classification (30,000-node taxonomy, three benchmarks, 9~LLMs across 4 families), we observe a regime shift from mandatory to opportunistic clarification, with Information-Seeking Effectiveness (ISE), a local diagnostic defined as the fraction of help interactions followed by a correct next navigation step (not a final-task metric), rising from 50% to 74%. Three diagnostic contrasts fail to reproduce this structure. A separability test shows that the information-seeking pattern (mode split, ISE ranking) persists when answer quality is degraded (-18.8% accuracy), supporting an empirical separation between where an agent seeks help and the quality of the help it receives. Under the controlled answer channel, accuracy gains reach +16.2% at 10-digit; we read this as an upper bound on what better localization could unlock, not a deployment estimate.
翻译:在分层推理中,失败通常源于中间决策点——智能体在未察觉自身缺乏关键信息时,便误入了错误分支。我们不将澄清视为外部不确定性触发机制,而是提出ACTION-RATING这一范式,将澄清操作置于智能体动作空间中,与导航动作共享同一有序量级,从而使"提问"在每个决策节点与"行动"直接竞争,并使寻求帮助的行为在中间状态变得可观测。根据智能体自身评分,自然涌现出两种结构迥异的信息寻求模式:强制性模式(无可行分支)与机会性模式(虽存在领先候选但仍有残余不确定性)。在协调关税表分类任务(30,000节点分类体系、三个基准测试、涵盖4个系列的9种大语言模型)中,我们观察到从强制性澄清向机会性澄清的范式转变,对应的信息寻求有效性(ISE,定义为后续正确导航步骤(非最终任务指标)之前帮助交互的比例)从50%提升至74%。三项诊断性对比实验均未能复现该结构。可分离性测试表明,当答案质量退化18.8%准确率时,信息寻求模式(模式分离、ISE排序)依然保持稳定,这支持了智能体寻求帮助的位置与其所获帮助质量之间的经验性分离。在受控答案通道下,10位编码精度的提升达16.2%;我们将其解读为更优定位能力可能解锁的上界,而非部署条件下的实际估计值。