A growing area of research investigates augmenting language models with tools (e.g., search engines, calculators) to overcome their shortcomings (e.g., missing or incorrect knowledge, incorrect logical inferences). Various few-shot tool-usage strategies have been proposed. However, there is no systematic and fair comparison across different strategies, or between these strategies and strong baselines that do not leverage tools. We conduct an extensive empirical analysis, finding that (1) across various datasets, example difficulty levels, and models, strong no-tool baselines are competitive to tool-assisted strategies, implying that effectively using tools with in-context demonstrations is a difficult unsolved problem; (2) for knowledge-retrieval tasks, strategies that *refine* incorrect outputs with tools outperform strategies that retrieve relevant information *ahead of* or *during generation*; (3) tool-assisted strategies are expensive in the number of tokens they require to work -- incurring additional costs by orders of magnitude -- which does not translate into significant improvement in performance. Overall, our findings suggest that few-shot tool integration is still an open challenge, emphasizing the need for comprehensive evaluations of future strategies to accurately assess their *benefits* and *costs*.
翻译:一个日益增长的研究领域探索通过增强语言模型与工具(如搜索引擎、计算器)的结合,以克服其局限性(例如知识缺失或错误、逻辑推理错误)。目前已提出多种少样本工具使用策略,然而,不同策略之间,以及这些策略与未借助工具的强基线方法之间,尚缺乏系统且公平的比较。我们开展了广泛的实证分析,发现:(1)在各类数据集、示例难度等级和模型上,未使用工具的强基线方法能与工具辅助策略相竞争,这表明利用上下文示例有效使用工具仍是一个尚未解决的难题;(2)在知识检索任务中,通过工具*修正*错误输出的策略,优于在*生成前*或*生成中*检索相关信息的策略;(3)工具辅助策略在所需令牌数量方面代价高昂——成本增加数个数量级——但这并未转化为性能的显著提升。总体而言,我们的发现表明,少样本工具集成仍是一个开放挑战,强调需对未来策略进行全面评估,以准确衡量其*收益*与*成本*。