Diffusion Transformers (DiTs) have become the dominant architecture for image and video generation, creating growing demand for efficient DiT serving. Existing systems assign each request a fixed parallel configuration throughout its lifetime. However, DiT workloads exhibit substantial heterogeneity across requests, execution stages, and system conditions, making static parallelism inefficient and often leading to poor GPU utilization and degraded service quality. This paper argues that DiT serving should treat GPU parallelism as a first-class schedulable resource. We present GF-DiT, a policy-programmable runtime for elastic DiT serving that dynamically adapts the parallelism of running requests according to workload demands and service objectives. GF-DiT introduces an asynchronous execution abstraction that decomposes requests into independently schedulable trajectory tasks and enables online GPU reallocation. To make elastic parallelism practical, GF-DiT further proposes group-free collectives, a lightweight communication abstraction that supports low-overhead online formation and reconfiguration of arbitrary execution groups. We implement GF-DiT in vLLM-Omni and evaluate it on representative image and video diffusion workloads. Compared with fixed-pipeline execution with static parallelism, GF-DiT improves throughput by up to 6.01$\times$, reduces mean latency by up to 95%, lowers SLO violation rates by up to 90%, and reduces communication-group setup overhead from 778 ms to approximately 60 $μ$s.


翻译:扩散变换器已成为图像与视频生成的主流架构,催生了对高效扩散变换器服务日益增长的需求。现有系统为每个请求在其生命周期内分配固定的并行配置。然而,扩散变换器工作负载在请求、执行阶段及系统条件间表现出显著异构性,导致静态并行策略效率低下,并常引发GPU利用率降低与服务质量下降。本文主张扩散变换器服务应将GPU并行视为一类可调度的首要资源。我们提出GF-DiT——一种策略可编程的弹性扩散变换器运行时系统,能根据工作负载需求与服务目标动态调整运行中请求的并行度。GF-DiT引入异步执行抽象机制,将请求分解为独立可调度的轨迹任务,并支持在线GPU重分配。为使弹性并行切实可行,GF-DiT进一步提出无组通信原语(group-free collectives),这是一种轻量级通信抽象,支持任意执行组的低开销在线组建与重构。我们在vLLM-Omni中实现GF-DiT,并在代表性图像与视频扩散工作负载上进行评估。与采用静态并行的固定流水线执行相比,GF-DiT将吞吐量提升高达6.01倍,平均延迟降低高达95%,服务等级协议违反率降低高达90%,并将通信组建开销从778毫秒降低至约60微秒。

0
下载
关闭预览

相关内容

中文版 | 集中式与分布式多智能体AI协调策略
专知会员服务
23+阅读 · 2025年5月8日
基于扩散模型和流模型的推理时引导生成技术
专知会员服务
17+阅读 · 2025年4月30日
面向图像处理逆问题的扩散模型研究综述
专知会员服务
16+阅读 · 2025年4月23日
Sora的幕后功臣?详解大火的DiT:拥抱Transformer的扩散模型
专知会员服务
42+阅读 · 2020年8月14日
盘点来自工业界的GPU共享方案
计算机视觉life
12+阅读 · 2021年9月2日
Deformable Kernels,用于图像/视频去噪,即将开源
极市平台
13+阅读 · 2019年8月29日
通用矩阵乘(GEMM)优化与卷积计算
极市平台
51+阅读 · 2019年6月19日
【干货】一文读懂什么是变分自编码器
专知
12+阅读 · 2018年2月11日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
VIP会员
最新内容
《无人机脆弱性利用:网络空间力量的新域》
专知会员服务
2+阅读 · 今天4:08
美空军如何将人工智能从战场部署至后方机关
专知会员服务
11+阅读 · 7月31日
《史诗怒火行动:多域前瞻评估》49页报告
专知会员服务
7+阅读 · 7月31日
《英国防部:未来空战系统数字化战略》33页
专知会员服务
5+阅读 · 7月31日
《面向自主飞行网络的智能体人工智能架构》
专知会员服务
7+阅读 · 7月31日
“史诗怒火”行动:现代多域作战的重要节点
专知会员服务
8+阅读 · 7月30日
《下一代无线网络中的多无人机通信资源管理》
相关基金
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
Top
微信扫码咨询专知VIP会员