Group Relative Policy Optimization has emerged as essential for aligning video diffusion models with human preferences, but faces a critical computational bottleneck: training a 14B parametered model typically demands hundreds of GPU days per experiment. Existing efficiency methods reduce costs through sliding window subsampling training timesteps, but fundamentally compromise optimization, exhibiting severe instability and failing to reach full trajectory performance. We present Flash-GRPO, a single-step training framework that outperforms full trajectory training in alignment quality under low computational budgets while substantially improving training efficiency. Flash-GRPO addresses two critical challenges: iso-temporal grouping eliminates timestep-confounded variance by enforcing prompt-wise temporal consistency, decoupling policy performance from timestep difficulty; temporal gradient rectification neutralizes the time-dependent scaling factor that causes vastly inconsistent gradient magnitudes across timesteps. Experiments on 1.3B to 14B parameter models validate Flash-GRPO's effectiveness, demonstrating substantial training acceleration with consistent stability and state-of-the-art alignment quality.
翻译:组相对策略优化已成为对齐视频扩散模型与人类偏好的关键方法,但面临严峻的计算瓶颈:训练一个140亿参数的模型通常需要每次实验数百个GPU天。现有效率方法通过滑动窗口子采样训练时间步来降低成本,但本质上损害了优化过程,表现出严重的不稳定性且无法达到完整轨迹性能。本文提出Flash-GRPO——一种单步训练框架,在低计算预算下其对齐质量超越完整轨迹训练,同时显著提升训练效率。Flash-GRPO解决两个关键挑战:等时分组通过强制按提示的时间一致性消除时间步混杂方差,将策略性能与时间步难度解耦;时间梯度校正中和了导致跨时间步梯度幅度严重不一致的时间相关缩放因子。在1.3B至14B参数模型上的实验验证了Flash-GRPO的有效性,表明其在保持稳定性和达到最先进对齐质量的同时实现了显著训练加速。