We study orchestration mechanisms for tool-using AI agents in realistic customer-service workflows over an unstructured knowledge base. We argue that declarative agents -- AI agents equipped with natural-language skill files appended to the system prompt -- are an effective orchestration paradigm. Concretely, we compare (i) a DeclarativeAgent that reads three domain-specific skill files at inference time and decides its own control flow, (ii) an ImperativeAgent based on a programmatic state machine with explicit phases, and (iii) an unscaffolded baseline agent modeled after the $τ$-Knowledge benchmark agent. Our ImperativeAgent is motivated by externalised-control inference as in Recursive Language Models and graph-based orchestration frameworks. We formalise the three agents as policy classes within a decentralised partially-observable Markov decision process and analyse their information-theoretic and structural properties; we then test the predicted differences empirically on five language models and two retrieval regimes. Our results show that retrieval quality is a dominant bottleneck for AI agents: when evidence is incomplete or skewed, all agents degrade substantially, and skill files cannot recover lost performance. Under high-quality retrieval, however, declarative skills consistently improve accuracy on procedural tasks and reduce orchestration errors, while the imperative state machine's brittleness does not reliably improve task success or compliance.
翻译:我们研究了在不具备结构化知识的现实客服工作流中,使用工具的AI代理的编排机制。我们认为声明式代理——即在系统提示中附加自然语言技能文件的AI代理——是一种有效的编排范式。具体而言,我们比较了:(i) 在推理时读取三个特定领域技能文件并自主决定控制流程的声明式代理(DeclarativeAgent);(ii) 基于具有显式阶段的程序化状态机的命令式代理(ImperativeAgent);以及 (iii) 以 $\tau$-Knowledge 基准代理为模型的无脚手架基线代理。我们的命令式代理源于递归语言模型和基于图的编排框架中的外部化控制推理。我们将这三种代理形式化为去中心化部分可观测马尔可夫决策过程中的策略类别,并分析其信息论和结构性属性;随后,我们在五种语言模型和两种检索机制下对预测的差异进行实证检验。结果表明,检索质量是AI代理的主要瓶颈:当证据不全或存在偏差时,所有代理的性能均显著下降,技能文件无法弥补性能损失。然而,在高检索质量条件下,声明式技能在程序性任务上持续提升准确性并减少编排错误,而命令式状态机的脆弱性并未可靠地改善任务成功率或合规性。