Statistical evaluation aims to estimate the generalization performance of a model using held-out i.i.d. test data sampled from the ground-truth distribution. In supervised learning settings such as classification, performance metrics such as error rate are well-defined, and test error reliably approximates population error given sufficiently large datasets. In contrast, evaluation is more challenging for generative models due to their open-ended nature: it is unclear which metrics are appropriate and whether such metrics can be reliably evaluated from finite samples. In this work, we introduce a theoretical framework for evaluating generative models and establish evaluability results for commonly used metrics. We study two categories of metrics: test-based metrics, including integral probability metrics (IPMs), and Rényi divergences. We show that IPMs with respect to any bounded test class can be evaluated from finite samples up to multiplicative and additive approximation errors. Moreover, when the test class has finite fat-shattering dimension, IPMs can be evaluated with arbitrary precision. In contrast, Rényi and KL divergences are not evaluable from finite samples, as their values can be critically determined by rare events. We also analyze the potential and limitations of perplexity as an evaluation method.
翻译:统计评估旨在利用从真实分布中采样的独立同分布测试数据估计模型的泛化性能。在分类等监督学习场景中,错误率等性能指标定义清晰,且给定足够大的数据集时,测试错误能够可靠地逼近总体错误。相比之下,由于生成模型的开放式特性,其评估更具挑战性:尚不明确哪些指标是合适的,以及这些指标能否基于有限样本进行可靠评估。本研究引入了一个评估生成模型的理论框架,并建立了常用指标的可评估性结果。我们研究两类指标:包括积分概率度量在内的基于测试的指标,以及Rényi散度。我们证明,对于任意有界测试类,积分概率度量可以通过有限样本进行乘法逼近误差和加法逼近误差的评估。此外,当测试类具有有限fat-shattering维度时,积分概率度量可以达到任意精度的评估。相比之下,Rényi散度和KL散度无法通过有限样本进行可评估,因为它们的取值可能由罕见事件决定性决定。我们还分析了困惑度作为评估方法的潜力与局限性。