This work proposes to measure the scope of a patent claim as the reciprocal of self-information contained in this claim. Self-information is calculated based on a probability of occurrence of the claim, where this probability is obtained from a language model. Grounded in information theory, this approach is based on the assumption that an unlikely concept is more informative than a usual concept, insofar as it is more surprising. In turn, the more surprising the information required to define the claim, the narrower its scope. Seven language models are considered, ranging from simplest models (each word or character has an identical probability) to intermediate models (based on average word or character frequencies), to large language models (LLMs) such as GPT2 and davinci-002. Remarkably, when using the simplest language models to compute the probabilities, the scope becomes proportional to the reciprocal of the number of words or characters involved in the claim, a metric already used in previous works. Application is made to multiple series of patent claims directed to distinct inventions, where each series consists of claims devised to have a gradually decreasing scope. The performance of the language models is then assessed through several ad hoc tests. The LLMs outperform models based on word and character frequencies, which themselves outdo the simplest models based on word or character counts. Interestingly, however, the character count appears to be a more reliable indicator than the word count.
翻译:本文提出将专利权利要求的范围度量定义为该权利要求所含自信息的倒数。自信息基于权利要求出现的概率计算,该概率由语言模型获得。基于信息论,该方法假设不常见概念比常见概念信息量更大,因其更具意外性。相应地,定义权利要求所需信息越令人意外,其范围越窄。研究考虑了七种语言模型,从最简单的模型(每个单词或字符具有相同概率)到中间模型(基于平均单词或字符频率),再到大型语言模型(如GPT2和davinci-002)。值得注意的是,使用最简单语言模型计算概率时,范围与权利要求所涉单词或字符数量的倒数成正比——这一度量指标已在先前研究中采用。将该方法应用于针对不同发明的多组专利权利要求序列,每组序列包含逐步缩小范围的权利要求。通过多项专项测试评估各语言模型性能:大型语言模型优于基于词频和字频的模型,而后者又优于基于单词或字符计数的简单模型。有趣的是,字符计数指标比单词计数指标更具可靠性。