Modern LLM serving is no longer homogeneous or monolithic. Production systems now combine disaggregated execution, complex parallelism, runtime optimizations, and stateful workloads such as reasoning, agents, and RL rollouts. Simulation is attractive for exploring this growing design space, yet existing simulators lack the architectural completeness and decision-grade fidelity it demands. Their monolithic-replica abstractions are ill-suited to disaggregated serving, while average-case analytical proxies can distort SLA predictions and even reverse optimization conclusions. We present Frontier, a discrete-event simulator for modern LLM inference serving. Frontier features a disaggregated abstraction. It captures the structure and dynamics of modern serving systems by modeling co-location, Prefill-Decode Disaggregation (PDD), and Attention-FFN Disaggregation (AFD) with role-specific cluster workers, incorporating key runtime optimizations (e.g., CUDA Graphs, speculative decoding) within the scheduler-batch-engine loop, and supporting stateful requests for emerging workloads. It further provides accurate and generalizable predictions of computation, communication, and memory costs across diverse serving scenarios with complex workload compositions. On 16-H800 GPU testbed, Frontier achieves an average throughput error below 4%. Compared with state-of-the-art simulators, it reduces end-to-end latency error from 44.9% to 6.4% under co-location and from 51.7% to 2.6% under disaggregation. It scales to over 1K GPUs on commodity CPUs and enables new use cases such as SLA-dependent Pareto frontier exploration, heterogeneous disaggregated allocation, agentic reasoning scheduling validation, and RL post-training reconfiguration.


翻译:现代LLM服务已不再是同质化或单一架构的。生产系统如今融合了分离式执行、复杂并行策略、运行时优化以及有状态工作负载(如推理、智能体和强化学习回滚)。模拟器在探索这一日益复杂的设计空间时极具吸引力,然而现有模拟器缺乏所需的架构完备性和决策级保真度。其单一副本抽象不适用于分离式服务,而平均情况的分析代理可能会扭曲服务水平协议(SLA)预测甚至导致优化结论反转。我们提出Frontier,一种面向现代LLM推理服务的离散事件模拟器。Frontier采用分离式抽象,通过建模协同部署、前缀-解码分离(PDD)和注意力-FFN分离(AFD)并引入角色特定的集群工作节点,捕捉现代服务系统的结构与动态;在调度-批处理-引擎循环中集成关键运行时优化(例如CUDA图形、推测解码);支持新兴工作负载的有状态请求。此外,它能在复杂工作负载组合的多样化服务场景下,提供对计算、通信和内存成本的精确且可泛化的预测。在16块H800 GPU测试平台上,Frontier的平均吞吐量误差低于4%。与最先进的模拟器相比,它将协同部署场景下的端到端延迟误差从44.9%降低至6.4%,将分离式场景下的误差从51.7%降低至2.6%。Frontier能够在商用CPU上扩展至超过1000块GPU,并支持新用例,如基于SLA的帕累托前沿探索、异构分离式分配、智能体推理调度验证以及强化学习后训练重配置。

0
下载
关闭预览

相关内容

《以人为中心的大型语言模型(LLM)研究综述》
专知会员服务
41+阅读 · 2024年11月25日
LLMCad:快速可扩展的设备上大型语言模型推理
专知会员服务
35+阅读 · 2023年9月11日
PlaNet 简介:用于强化学习的深度规划网络
谷歌开发者
13+阅读 · 2019年3月16日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
8+阅读 · 2015年12月31日
VIP会员
最新内容
俄乌无人机战争的六大启示
专知会员服务
9+阅读 · 8月3日
《无人机空中监控:通信实验洞察》
专知会员服务
7+阅读 · 8月3日
从采集到决策:美军视角下的战术情报范式重构
《履带式无人地面战车技术发展现状》
专知会员服务
7+阅读 · 8月2日
《无人机脆弱性利用:网络空间力量的新域》
专知会员服务
10+阅读 · 8月1日
相关基金
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
8+阅读 · 2015年12月31日
Top
微信扫码咨询专知VIP会员