We discuss probabilistic neural network models for unsupervised learning where the distribution of the hidden layer is fixed. We argue that learning machines with this architecture enjoy a number of desirable properties. For example, the model can be chosen as a simple and interpretable one, it does not need to be over-parametrised and training is argued to be efficient in a thermodynamic sense. When hidden units are binary variables, these models have a natural interpretation in terms of features. We show that the featureless state corresponds to a state of maximal ignorance about the features and that learning the first feature depends on non-Gaussian statistical properties of the data. We suggest that the distribution of hidden variables should be chosen according to the principle of maximal relevance. We introduce the Hierarchical Feature Model (HFM) as an example of a model that satisfies this principle, and that encodes a neutral a priori organisation of the feature space. We present extensive numerical experiments in order i) to test that the internal representation of learning machines can indeed be independent of the data with which they are trained and ii) that only a finite number of features are needed to describe a number of datasets.
翻译:我们讨论用于无监督学习的概率神经网络模型,其中隐藏层的分布是固定的。我们认为,具有这种架构的学习机器具有许多理想特性。例如,该模型可以选择为简单且可解释的模型,无需过度参数化,并且在热力学意义上训练是高效的。当隐藏单元为二元变量时,这些模型在特征方面具有自然解释。我们表明,无特征状态对应于对特征最大无知的状态,而学习第一个特征依赖于数据的非高斯统计特性。我们建议应根据最大相关性原则选择隐藏变量的分布。我们引入层次特征模型(HFM)作为满足该原则的示例,该模型编码了特征空间的先验中性组织。我们进行大量数值实验,以i)测试学习机器的内部表示确实可以独立于其训练数据,以及ii)描述多个数据集仅需有限数量的特征。