This paper explores techniques that focus on understanding and resolving ambiguity in language within the field of natural language processing (NLP), highlighting the complexity of linguistic phenomena such as polysemy and homonymy and their implications for computational models. Focusing extensively on Word Sense Disambiguation (WSD), it outlines diverse approaches ranging from deep learning techniques to leveraging lexical resources and knowledge graphs like WordNet. The paper introduces cutting-edge methodologies like word sense extension (WSE) and neuromyotonic approaches, enhancing disambiguation accuracy by predicting new word senses. It examines specific applications in biomedical disambiguation and language specific optimisation and discusses the significance of cognitive metaphors in discourse analysis. The research identifies persistent challenges in the field, such as the scarcity of sense annotated corpora and the complexity of informal clinical texts. It concludes by suggesting future directions, including using large language models, visual WSD, and multilingual WSD systems, emphasising the ongoing evolution in addressing lexical complexities in NLP. This thinking perspective highlights the advancement in this field to enable computers to understand language more accurately.
翻译:本文探讨了自然语言处理领域中关注理解与消解语言歧义的技术,重点阐述了多义性、同形异义等语言现象的复杂性及其对计算模型的影响。研究聚焦于词义消歧,系统梳理了从深度学习技术到利用词汇资源及WordNet等知识图谱的多种方法。本文介绍了词义扩展与神经肌张力方法等前沿技术,通过预测新词义提升消歧精度。研究考察了生物医学消歧与语言特定优化等具体应用场景,并讨论了认知隐喻在话语分析中的重要性。本文识别了该领域持续存在的挑战,如标注语料库稀缺及非正式临床文本的复杂性。最后指出未来研究方向,包括利用大语言模型、视觉词义消歧和多语言词义消歧系统,强调自然语言处理领域应对词汇复杂性的持续演进。这一思考视角凸显了该领域使计算机更精准理解语言的前沿进展。