Herding is a deterministic algorithm used to generate data points that can be regarded as random samples satisfying input moment conditions. The algorithm is based on the complex behavior of a high-dimensional dynamical system and is inspired by the maximum entropy principle of statistical inference. In this paper, we propose an extension of the herding algorithm, called entropic herding, which generates a sequence of distributions instead of points. Entropic herding is derived as the optimization of the target function obtained from the maximum entropy principle. Using the proposed entropic herding algorithm as a framework, we discuss a closer connection between herding and the maximum entropy principle. Specifically, we interpret the original herding algorithm as a tractable version of entropic herding, the ideal output distribution of which is mathematically represented. We further discuss how the complex behavior of the herding algorithm contributes to optimization. We argue that the proposed entropic herding algorithm extends the application of herding to probabilistic modeling. In contrast to original herding, entropic herding can generate a smooth distribution such that both efficient probability density calculation and sample generation become possible. To demonstrate the viability of these arguments in this study, numerical experiments were conducted, including a comparison with other conventional methods, on both synthetic and real data.
翻译:聚集是一种确定性算法,用于生成可视为满足输入矩条件的随机样本的数据点。该算法基于高维动力系统的复杂行为,并受统计推断中最大熵原理的启发。本文提出了一种称为熵驱动聚集的聚集算法扩展,它生成一系列分布而非点序列。熵驱动聚集源自对从最大熵原理获得的目标函数的优化。利用所提出的熵驱动聚集算法框架,我们讨论了聚集算法与最大熵原理之间更紧密的联系。具体而言,我们将原始聚集算法解释为熵驱动聚集的一个易处理版本,其理想输出分布在数学上得以表征。我们进一步探讨了聚集算法的复杂行为如何促进优化。研究表明,所提出的熵驱动聚集算法将聚集的应用扩展到概率建模领域。与原始聚集不同,熵驱动聚集能生成平滑分布,从而同时实现高效的概率密度计算与样本生成。为验证本文的这些论点,我们在合成数据与真实数据上进行了数值实验,包括与其他传统方法的对比分析。