We present a computationally efficient framework, called $\texttt{FlowDRO}$, for solving flow-based distributionally robust optimization (DRO) problems with Wasserstein uncertainty sets while aiming to find continuous worst-case distribution (also called the Least Favorable Distribution, LFD) and sample from it. The requirement for LFD to be continuous is so that the algorithm can be scalable to problems with larger sample sizes and achieve better generalization capability for the induced robust algorithms. To tackle the computationally challenging infinitely dimensional optimization problem, we leverage flow-based models and continuous-time invertible transport maps between the data distribution and the target distribution and develop a Wasserstein proximal gradient flow type algorithm. In theory, we establish the equivalence of the solution by optimal transport map to the original formulation, as well as the dual form of the problem through Wasserstein calculus and Brenier theorem. In practice, we parameterize the transport maps by a sequence of neural networks progressively trained in blocks by gradient descent. We demonstrate its usage in adversarial learning, distributionally robust hypothesis testing, and a new mechanism for data-driven distribution perturbation differential privacy, where the proposed method gives strong empirical performance on high-dimensional real data.
翻译:我们提出了一种计算高效的框架,称为$\texttt{FlowDRO}$,用于解决基于Wasserstein不确定集、同时旨在找到连续最坏情况分布(也称为最不利分布,LFD)并从中采样的流式分布鲁棒优化(DRO)问题。要求LFD具有连续性,是为了使算法能够扩展到更大样本量的问题,并提升所诱导的鲁棒算法的泛化能力。为应对该计算上极具挑战性的无穷维优化问题,我们利用基于流的模型以及数据分布与目标分布之间的连续时间可逆传输映射,发展了一种Wasserstein近端梯度流类型算法。理论上,我们通过最优传输映射建立了原始问题解与原始形式的等价性,并利用Wasserstein微积分和Brenier定理推导了问题的对偶形式。实践中,我们通过一系列神经网络以梯度下降方式逐块渐进训练来参数化传输映射。我们在对抗学习、分布鲁棒假设检验以及一种基于数据驱动分布扰动差分隐私的新机制中展示了该方法的有效性,所提方法在高维真实数据上展现了强大的实证性能。