$f \propto r^{-\alpha} \cdot (r+\gamma)^{-\beta}$ has been empirically shown more precise than a na\"ive power law $f\propto r^{-\alpha}$ to model the rank-frequency ($r$-$f$) relation of words in natural languages. This work shows that the only crucial parameter in the formulation is $\gamma$, which depicts the resistance to vocabulary growth on a corpus. A method of parameter estimation by searching an optimal $\gamma$ is proposed, where a ``zeroth word'' is introduced technically for the calculation. The formulation and parameters are further discussed with several case studies.
翻译:$f \propto r^{-\alpha} \cdot (r+\gamma)^{-\beta}$已被经验证明比朴素幂律$f\propto r^{-\alpha}$更能精确建模自然语言词汇的等级-频率($r$-$f$)关系。本研究表明,该公式中唯一的关键参数是$\gamma$,它表征了语料库词汇增长受到的抵抗。本文提出了一种通过搜索最优$\gamma$进行参数估计的方法,其中技术上引入了“第零词汇”用于计算。通过多个案例研究,进一步讨论了该公式及其参数。