Word sense disambiguation primarily addresses the lexical ambiguity of common words based on a predefined sense inventory. Conversely, proper names are usually considered to denote an ad-hoc real-world referent. Once the reference is decided, the ambiguity is purportedly resolved. However, proper names also exhibit ambiguities through appellativization, i.e., they act like common words and may denote different aspects of their referents. We proposed to address the ambiguities of proper names through the light of regular polysemy, which we formalized as dot objects. This paper introduces a combined word sense disambiguation (WSD) model for disambiguating common words against Chinese Wordnet (CWN) and proper names as dot objects. The model leverages the flexibility of a gloss-based model architecture, which takes advantage of the glosses and example sentences of CWN. We show that the model achieves competitive results on both common and proper nouns, even on a relatively sparse sense dataset. Aside from being a performant WSD tool, the model further facilitates the future development of the lexical resource.
翻译:词义消歧主要基于预定义词义清单处理普通词汇的词汇歧义。相反,专有名词通常被认为指代特定现实世界中的实体,一旦指代对象确定,其歧义性便得以化解。然而,专有名词也通过名词化呈现歧义——即它们像普通词汇一样运作,可能指代其指称对象的不同方面。我们提出通过常规多义现象(将其形式化为点对象)来处理专有名词的歧义问题。本文介绍了一种结合词义消歧(WSD)的混合模型,该模型可针对中文词网(CWN)的普通词汇与作为点对象的专有名词进行消歧。该模型利用基于释义的灵活架构,充分运用CWN的释义和例句。研究表明,即使在相对稀疏的词义数据集上,该模型在普通名词和专有名词上均取得了具有竞争力的结果。除了作为高效WSD工具外,该模型还进一步推动了词汇资源的未来开发。