The article introduces corrections to Zipf's and Heaps' laws based on systematic models of the hapax rate. The derivation rests on two assumptions: The first one is the standard urn model which predicts that marginal frequency distributions for shorter texts look as if word tokens were sampled blindly from a given longer text. The second assumption posits that the rate of hapaxes is a simple function of the text size. Four such functions are discussed: the constant model, the Davis model, the linear model, and the logistic model. It is shown that the logistic model yields the best fit.
翻译:本文基于罕用词率的系统性模型,提出了对齐普夫定律和希普斯定律的修正。推导过程基于两个假设:第一个假设是标准瓮模型,该模型预测较短文本的边缘频率分布,如同单词标记是从给定较长文本中随机抽取所得。第二个假设假定罕用词率是文本规模的简单函数。本文讨论了四种此类函数:常数模型、戴维斯模型、线性模型和逻辑斯蒂模型。研究表明,逻辑斯蒂模型拟合效果最佳。