Deploying LLM agents at scale typically requires choosing between quality and cost. Existing cost-reduction approaches fail to preserve agility: the ability to iterate rapidly without human time bottlenecks. Prompt engineering is brittle and slows iteration, while fine-tuning requires multi-day training and commitment to fixed designs; both are impractical for iterative workflows and time-sensitive batch jobs. We demonstrate that established inference-time techniques--dynamic in-context learning and self-consistency cascades--can be leveraged to shift the cost-accuracy Pareto frontier while preserving agility. Practitioners run the teacher on a small task subset to collect demonstrations, then immediately deploy a cheaper student on the remainder. At each step, the system retrieves relevant teacher demonstrations as in-context examples. When multiple student samples agree, we proceed; when they diverge, we fall back to the teacher. This requires no prompt engineering or training. On ALFWorld, we match teacher accuracy at 2.5x lower cost (0.059 to 0.024 per episode). On AppWorld, we achieve 3.5x cost reduction while recovering 79% of teacher accuracy. Our empirical analyses provide guidance on key design choices: teacher database size, demonstration set size, retrieval strategy, and cascade thresholds. These analyses highlight inference-time levers for navigating cost-performance tradeoffs without sacrificing human development speed.


翻译:大规模部署LLM智能体通常需要在质量与成本之间做出权衡。现有降本方法无法保持敏捷性——即无需人类时间瓶颈即可快速迭代的能力。提示工程脆弱且拖慢迭代速度,而微调则需要多天训练并固守固定设计;两者对于迭代工作流和时间敏感的批量任务均不切实际。我们证明,成熟的推理时技术——动态上下文学习和自一致性级联——可用于移动成本-准确率帕累托前沿,同时保持敏捷性。实践者可在小型任务子集上运行教师模型以收集演示样本,随后立即在剩余任务上部署更廉价的学生模型。系统在每一步检索相关教师演示作为上下文示例:当多个学生样本一致时继续推进,出现分歧时则回退至教师模型。这无需提示工程或训练。在ALFWorld上,我们以2.5倍成本降低(每轮0.059 vs 0.024美元)达到教师准确率。在AppWorld上,我们实现3.5倍成本降低并恢复79%的教师准确率。实证分析为关键设计选择提供了指导:教师数据库规模、演示集大小、检索策略及级联阈值。这些分析揭示了无需牺牲人类开发速度即可导航成本-性能权衡的推理时杠杆。

0
下载
关闭预览

相关内容

智能体工程(Agent Engineering)
专知会员服务
39+阅读 · 2025年12月31日
高效大语言模型推理服务综述
专知会员服务
18+阅读 · 2025年4月30日
AI Agent,大模型时代重要落地方向, 42页ppt
专知会员服务
292+阅读 · 2023年10月12日
探索(Exploration)还是利用(Exploitation)?强化学习如何tradeoff?
深度强化学习实验室
13+阅读 · 2020年8月23日
最全的智慧工地解决方案
智能交通技术
11+阅读 · 2019年8月30日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
VIP会员
最新内容
边缘计算的军事应用
专知会员服务
5+阅读 · 8月9日
一种考虑资源机动性的武器目标分配混合算法
专知会员服务
8+阅读 · 8月8日
《多域冲突比较支持模型》60页
专知会员服务
13+阅读 · 8月7日
相关基金
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
4+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员