We show that Fréchet Distance (FD), long considered impractical as a training objective, can in fact be effectively optimized in the representation space. Our idea is simple: decouple the population size for FD estimation (e.g., 50k) from the batch size for gradient computation (e.g., 1024). We term this approach FD-loss. Optimizing FD-loss reveals several surprising findings. First, post-training a base generator with FD-loss in different representation spaces consistently improves visual quality. Under the Inception feature space, a one-step generator achieves0.72 FID on ImageNet 256x256. Second, the same FD-loss repurposes multi-step generators into strong one-step generators without teacher distillation, adversarial training or per-sample targets. Third, FID can misrank visual quality: modern representations can yield better samples despite worse Inception FID. This motivates FDr$^k$, a multi-representation metric. We hope this work will encourage further exploration of distributional distances in diverse representation spaces as both training objectives and evaluation metrics for generative models.
翻译:我们表明,长期被认为不适合作为训练目标的弗雷歇距离(FD),实际上可以在表示空间中被有效优化。我们的想法很简单:将FD估计的总体规模(例如50k)与梯度计算的批次规模(例如1024)解耦。我们将此方法称为FD-loss。优化FD-loss揭示了几个令人惊讶的发现。首先,在不同表示空间中使用FD-loss对基础生成器进行后训练,能持续提升视觉质量。在Inception特征空间下,一个一步生成器在ImageNet 256x256上达到了0.72的FID。其次,相同的FD-loss可以将多步生成器转化为强大的单步生成器,无需教师蒸馏、对抗训练或逐样本目标。第三,FID可能错误排序视觉质量:尽管Inception FID更差,现代表示可能产生更好的样本。这催生了FDr$^k$,一种多表示度量标准。我们希望这项工作能鼓励在多样化的表示空间中进一步探索分布距离,既作为生成模型的训练目标,也作为评估指标。