The power law is useful in describing count phenomena such as network degrees and word frequencies. With a single parameter, it captures the main feature that the frequencies are linear on the log-log scale. Nevertheless, there have been criticisms of the power law, for example that a threshold needs to be pre-selected without its uncertainty quantified, that the power law is simply inadequate, and that subsequent hypothesis tests are required to determine whether the data could have come from the power law. We propose a modelling framework that combines two different generalisations of the power law, namely the generalised Pareto distribution and the Zipf-polylog distribution, to resolve these issues. The proposed mixture distributions are shown to fit the data well and quantify the threshold uncertainty in a natural way. A model selection step embedded in the Bayesian inference algorithm further answers the question whether the power law is adequate.
翻译:幂律在描述网络度数和词频等计数现象时非常有用。通过单一参数,它能捕捉到在对数-对数尺度上频率呈线性这一主要特征。然而,幂律也受到了一些批评,例如需要预先选择一个阈值而不量化其不确定性、幂律本身可能不够充分,以及需要后续假设检验来确定数据是否源自幂律。为解决这些问题,我们提出一个结合了幂律的两种不同推广形式——广义帕累托分布和齐普夫-多对数分布——的建模框架。所提出的混合分布能够很好地拟合数据,并以自然的方式量化阈值不确定性。嵌入在贝叶斯推断算法中的模型选择步骤进一步回答了幂律是否充分的问题。