We propose a notion of common information that allows one to quantify and separate the information that is shared between two random variables from the information that is unique to each. Our notion of common information is defined by an optimization problem over a family of functions and recovers the G\'acs-K\"orner common information as a special case. Importantly, our notion can be approximated empirically using samples from the underlying data distribution. We then provide a method to partition and quantify the common and unique information using a simple modification of a traditional variational auto-encoder. Empirically, we demonstrate that our formulation allows us to learn semantically meaningful common and unique factors of variation even on high-dimensional data such as images and videos. Moreover, on datasets where ground-truth latent factors are known, we show that we can accurately quantify the common information between the random variables.
翻译:我们提出了一种公共信息的概念,该概念能够量化并分离两个随机变量之间共享的信息与各自独有的信息。我们的公共信息定义基于一族函数的优化问题,并作为特例恢复了Gács-Körner公共信息。重要的是,该概念可通过底层数据分布的样本进行经验近似。随后,我们提供了一种方法,通过对传统变分自编码器进行简单修改,来划分并量化公共信息与独有信息。实验表明,即使在图像和视频等高维数据上,我们的方法也能学习到具有语义意义的公共与独有变异因子。此外,在已知真实隐因子的数据集上,我们证明该方法能够准确量化随机变量之间的公共信息。