Two widely adopted techniques for LLM inference serving systems today are hybrid batching and disaggregated serving. A hybrid batch combines prefill and decode tokens of different requests in the same batch to improve resource utilization and throughput at the cost of increased latency per token. In contrast, disaggregated serving decouples compute-bound prefill and bandwidth-bound decode phases to optimize for service level objectives (SLOs) at the cost of resource under-utilization and KV-cache transfer overheads. To address the limitations of these techniques, we propose RAPID-Serve: a technique to concurrently execute prefill and decode on the same GPU(s) to meet latency SLOs while maintaining high throughput and efficient resource utilization. Furthermore, we propose Adaptive Resource Management for runtime compute resource allocation, optionally leveraging CU masking (a fine-grained Compute Unit partitioning feature on AMD Instinct\textsuperscript{TM} GPUs). RAPID-Serve provides up to 4.1x (average 1.7x) unconstrained throughput improvement and 32x and higher (average 4.9x) throughput improvement under SLO constraints, showing it as an effective strategy compared to the state-of-the-art approaches, particularly in resource-constrained environments.


翻译:当前,大语言模型推理服务系统广泛采用的两项技术是混合批处理与解耦服务。混合批处理将不同请求的预填充和解码令牌组合在同一批次中,以提高资源利用率和吞吐量,但代价是增加了每个令牌的延迟。相比之下,解耦服务将计算密集的预填充阶段与带宽密集的解码阶段解耦,以优化服务级别目标,但代价是资源利用不足和KV缓存传输开销。为克服这些技术的局限性,我们提出RAPID-Serve:一种在同一GPU上并发执行预填充和解码的技术,旨在满足延迟SLO的同时,保持高吞吐量和高效的资源利用率。此外,我们提出了自适应资源管理方案,用于运行时计算资源分配,并可选择性地利用CU掩码(AMD Instinct\textsuperscript{TM} GPU上的一种细粒度计算单元分区功能)。RAPID-Serve在无约束条件下可实现高达4.1倍(平均1.7倍)的吞吐量提升,在SLO约束下可实现32倍及以上(平均4.9倍)的吞吐量提升,这表明与现有先进方法相比,尤其是在资源受限的环境中,它是一种有效的策略。

0
下载
关闭预览

相关内容

TensorFlow 2.0新特性之Ragged Tensor
深度学习每日摘要
18+阅读 · 2019年4月5日
国家自然科学基金
0+阅读 · 2017年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
VIP会员
最新内容
ICML 2026 教程 | 数值优化理论还重要吗?
专知会员服务
1+阅读 · 今天11:09
ICM 2026 | 陶哲轩:人工智能时代的数学
专知会员服务
0+阅读 · 今天11:05
《反无人机交战场景下的战斗归零研究》
专知会员服务
2+阅读 · 今天2:34
博士论文 | 用代码结构感知方法推进代码大模型
《决策模型比较研究》
专知会员服务
11+阅读 · 7月25日
《美军水下战与海床战概述及本地实施》
专知会员服务
6+阅读 · 7月25日
面向未来冲突推进陆军情报体制改革
专知会员服务
5+阅读 · 7月25日
相关VIP内容
相关基金
国家自然科学基金
0+阅读 · 2017年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员