Generative models are increasingly deployed as substitutes for real data in downstream scientific workflows, yet standard evaluation criteria remain focused on marginal distribution matching. We argue that this represents a fundamental gap: downstream inference is rarely a marginal operation, and a model that passes every univariate diagnostic can still produce structurally unreliable synthetic data. We introduce covariance-level dependence fidelity, measured by D_Sigma(P,Q) = ||Sigma_P - Sigma_Q||_F, as a principled, computable criterion for evaluating whether a generative model preserves the joint structure of data beyond its univariate marginals. Three results formalise this criterion. First, marginal fidelity provides no constraint on dependence structure: D_Sigma can be made arbitrarily large while all univariate marginals match exactly. Second, covariance divergence induces quantifiable downstream instability, including sign reversals in population regression coefficients. Third, bounding D_Sigma provides positive stability guarantees for dependence-sensitive procedures such as PCA via Davis-Kahan-type bounds. Empirical validation across three domains, image data (Fashion-MNIST VAE, n = 60,000), bulk RNA-seq (TCGA-BRCA, n = 1,111), and a small-sample stress test (Alzheimer's gene expression, n = 113), shows that D_Sigma/delta consistently distinguishes structure-discarding from structure-preserving generators in cases where standard marginal diagnostics show little separation, confirming that covariance-level fidelity provides information orthogonal to existing evaluation metrics across domains and sample sizes.
翻译:生成模型越来越多地被部署为下游科学工作流中真实数据的替代品,然而标准的评估标准仍集中于边际分布匹配。我们认为这代表了一个根本性缺陷:下游推理很少是边际操作,一个通过了所有单变量诊断的模型仍然可能产生结构不可靠的合成数据。我们引入了协方差层面的依赖保真度,通过 D_Sigma(P,Q) = ||Sigma_P - Sigma_Q||_F 衡量,将其作为一个原则性且可计算的标准,用于评估生成模型是否保留了超越单变量边际的数据联合结构。三个结果将这一标准形式化。首先,边际保真度对依赖结构没有约束:当所有单变量边际完全匹配时,D_Sigma 可以任意大。其次,协方差散度会导致可量化的下游不稳定性,包括总体回归系数的符号反转。第三,限制 D_Sigma 可为依赖敏感的程序(例如主成分分析,通过 Davis-Kahan 类界)提供正向的稳定性保证。在三个领域(图像数据(Fashion-MNIST VAE,n = 60,000)、批量 RNA-seq 数据(TCGA-BRCA,n = 1,111)以及小样本压力测试(阿尔茨海默症基因表达数据,n = 113))的实证验证表明,在标准边际诊断显示区分度很小的情况下,D_Sigma/delta 能一致地区分丢弃结构与保留结构的生成器,证实了协方差层面的保真度在不同领域和样本量下提供了与现有评估指标正交的信息。