Product of experts (PoE) are layered networks in which the value at each node is an AND (or product) of the values (possibly negated) at its inputs. These were introduced as a neural network architecture that can efficiently learn to generate high-dimensional data which satisfy many low-dimensional constraints -- thereby allowing each individual expert to perform a simple task. PoEs have found a variety of applications in learning. We study the problem of identifiability of a product of experts model having a layer of binary latent variables, and a layer of binary observables that are iid conditional on the latents. The previous best upper bound on the number of observables needed to identify the model was exponential in the number of parameters. We show: (a) When the latents are uniformly distributed, the model is identifiable with a number of observables equal to the number of parameters (and hence best possible). (b) In the more general case of arbitrarily distributed latents, the model is identifiable for a number of observables that is still linear in the number of parameters (and within a factor of two of best-possible). The proofs rely on root interlacing phenomena for some special three-term recurrences.
翻译:专家乘积模型(Product of Experts, PoE)是一种分层网络,其中每个节点的值是其输入值(可能取反)的AND(或乘积)。这些模型最初被提出作为一种神经网络架构,能够高效学习生成满足多个低维约束的高维数据——从而允许每个专家独立执行简单任务。PoE在学习领域已有多种应用。我们研究具有二元潜变量层和以潜变量为条件的独立同分布二元观测变量层的专家乘积模型的可辨识性问题。此前识别该模型所需观测数量的最佳上界是参数数量的指数级。我们证明:(a) 当潜变量均匀分布时,模型在观测数量等于参数数量时可辨识(即达到理论最优);(b) 在潜变量任意分布的一般情况下,模型在观测数量仍为参数数量的线性阶(并在两倍因子内接近最优)时可辨识。证明依赖于某些特殊三项递推关系的根交错现象。