Prefill-decode (PD) disaggregation has become the standard architecture for large-scale LLM serving, but in practice its deployment boundary is still determined by KVCache transfer. In conventional dense-attention models, prefill generates huge KVCache traffics that keep prefill and decode tightly coupled within a single high-bandwidth network domain, limiting heterogeneous deployment and resource elasticity. Recent hybrid-attention architectures substantially reduce KVCache size, making cross-cluster KVCache transport increasingly plausible. However, smaller KVCache alone does not make heterogeneous cross-datacenter PD serving practical: real workloads remain bursty, request lengths are highly skewed, prefix caches are unevenly distributed, and inter-cluster bandwidth fluctuates. A naive design that fully externalizes prefill can therefore still suffer from congestion, unstable queueing, and poor utilization. We present Prefill-as-a-Service (PrfaaS), a cross-datacenter serving architecture that selectively offloads long-context prefill to standalone, compute-dense prefill clusters and transfers the resulting KVCache over commodity Ethernet to local PD clusters for decode. Rather than treating reduced KVCache as sufficient, PrfaaS combines model-side KV efficiency with system-side selective offloading, bandwidth-aware scheduling, and cache-aware request placement. This design removes the requirement that heterogeneous accelerators share the same low-latency RDMA fabric, enabling independent scaling of prefill and decode capacity across loosely coupled clusters. In a case study using an internal 1T-parameter hybrid model, a PrfaaS-augmented heterogeneous deployment achieves 54% and 32% higher serving throughput than homogeneous PD and naive heterogeneous baselines, respectively, while consuming only modest cross-datacenter bandwidth.


翻译:预填充-解码(PD)分离已成为大规模LLM服务部署的标准架构,但在实际应用中其部署边界仍受限于KVCache传输。在传统密集注意力模型中,预填充阶段会产生大量KVCache流量,迫使预填充与解码紧密耦合于单一高带宽网络域内,限制了异构部署和资源弹性。近期提出的混合注意力架构显著缩减了KVCache体积,使得跨集群KVCache传输逐渐具备可行性。然而,仅凭更小的KVCache不足以使异构跨数据中心PD服务落地:实际工作负载仍具有突发性,请求长度高度偏斜,前缀缓存分布不均,集群间带宽存在波动。因此,若简单将预填充完全外置化,仍会面临网络拥塞、排队不稳定及资源利用率低下的问题。本文提出"预填充即服务"(PrfaaS),一种跨数据中心服务架构:该架构将长上下文预填充任务选择性卸载至独立的计算密集型预填充集群,并通过以太网将生成的KVCache传输至本地PD集群进行解码。PrfaaS并未将KVCache缩减视为充分条件,而是将模型层面的KV效率与系统层面的选择性卸载、带宽感知调度及缓存感知请求放置策略相结合。该设计消除了异构加速器必须共享同一低延迟RDMA网络的硬性约束,使松散耦合集群中的预填充与解码能力可独立扩展。基于某内部1T参数混合模型的案例研究显示,采用PrfaaS增强的异构部署相比同构PD基线方案提升服务吞吐量54%,相比朴素异构基线方案提升32%,且仅消耗适度的跨数据中心带宽。

0
下载
关闭预览

相关内容

TransMLA:多头潜在注意力(MLA)即为所需
专知会员服务
23+阅读 · 2025年2月13日
面向多模态智能的下一个Token预测:综述
专知会员服务
26+阅读 · 2024年12月30日
并行算法演进,从MapReduce到MPI
凡人机器学习
10+阅读 · 2017年11月5日
国家自然科学基金
0+阅读 · 2017年12月31日
国家自然科学基金
1+阅读 · 2017年12月31日
国家自然科学基金
2+阅读 · 2017年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
9+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
2+阅读 · 2014年12月31日
VIP会员
最新内容
《无人机对海面作战影响评估》
专知会员服务
10+阅读 · 7月21日
印度精确打击与指挥架构的断层
专知会员服务
5+阅读 · 7月20日
美空军AI完成F-16战斗机自主空战历史性试飞
专知会员服务
6+阅读 · 7月20日
深入Project Maven:为何人工智能在战场上依然失灵
锻造未来士兵:外骨骼、基因工程与赛博格
专知会员服务
8+阅读 · 7月19日
相关VIP内容
相关基金
国家自然科学基金
0+阅读 · 2017年12月31日
国家自然科学基金
1+阅读 · 2017年12月31日
国家自然科学基金
2+阅读 · 2017年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
9+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
2+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员