With the increasing demand of intelligent systems capable of operating in different user contexts (e.g. users on the move) the correct interpretation of the user-need by such systems has become crucial to give a consistent answer to the user query. The most effective techniques which are used to address such task are in the fields of natural language processing and semantic expansion of terms. Such systems are aimed at estimating the actual meaning of input queries, addressing the concepts of the words which are expressed within the user questions. The aim of this paper is to demonstrate which semantic relation impacts the most in semantic expansion-based retrieval systems and to identify the best tradeoff between accuracy and noise introduction when combining such relations. The evaluations are made building a simple natural language processing system capable of querying any taxonomy-driven domain, making use of the combination of different semantic expansions as knowledge resources. The proposed evaluation employs a wide and varied taxonomy as a use-case, exploiting its labels as basis for the expansions. To build the knowledge resources several corpora have been produced and integrated as gazetteers into the NLP infrastructure with the purpose of estimating the pseudo-queries corresponding to the taxonomy labels, considered as the possible intents.
翻译:随着能够在不同用户情境(例如移动中的用户)下运行的智能系统需求日益增长,正确解读用户需求已成为系统对用户查询提供一致响应的关键。解决该任务最有效的技术属于自然语言处理和术语语义扩展领域。此类系统旨在估计输入查询的实际含义,处理用户问题中表达的词项概念。本文旨在论证语义扩展式检索系统中何种语义关系影响最大,并确定组合此类关系时准确性与引入噪声之间的最佳平衡。评估通过构建一个能够查询任意分类驱动领域的简易自然语言处理系统实现,该系统利用不同语义扩展组合作为知识资源。所提出的评估采用宽泛且多样的分类体系作为用例,利用其标签作为扩展基础。为构建知识资源,已生成并集成多个语料库作为NLP基础设施中的地名辞典,旨在估计与分类标签相对应的伪查询(视为可能的意图)。