Bayesian Pseudo-Coreset (BPC) and Dataset Condensation are two parallel streams of work that construct a synthetic set such that, a model trained independently on this synthetic set, yields the same performance as training on the original training set. While dataset condensation methods use non-bayesian, heuristic ways to construct such a synthetic set, BPC methods take a bayesian approach and formulate the problem as divergence minimization between posteriors associated with original data and synthetic data. However, BPC methods generally rely on distributional assumptions on these posteriors which makes them less flexible and hinders their performance. In this work, we propose to solve these issues by modeling the posterior associated with synthetic data by an energy-based distribution. We derive a contrastive-divergence-like loss function to learn the synthetic set and show a simple and efficient way to estimate this loss. Further, we perform rigorous experiments pertaining to the proposed method. Our experiments on multiple datasets show that the proposed method not only outperforms previous BPC methods but also gives performance comparable to dataset condensation counterparts.
翻译:贝叶斯伪核心集(BPC)与数据集压缩是两条并行的研究路径,均旨在构建一个合成数据集,使得在该合成数据集上独立训练的模型,能够产生与在原始训练集上训练相同的性能。数据集压缩方法采用非贝叶斯、启发式的方式构建此类合成集,而BPC方法则基于贝叶斯视角,将问题形式化为原始数据与合成数据对应后验之间的散度最小化。然而,BPC方法通常依赖这些后验的分布假设,这降低了其灵活性并阻碍了性能提升。在本工作中,我们提出通过能量基分布对合成数据对应的后验进行建模以解决上述问题。我们推导出一种类对比散度的损失函数来学习合成集,并展示了该损失函数的简单高效估计方法。此外,我们围绕所提方法进行了严格的实验。在多个数据集上的实验表明,所提方法不仅优于先前的BPC方法,并且性能与数据集压缩方法相当。