With the advent of technology and use of latest devices, they produces voluminous data. Out of it, 80% of the data are unstructured and remaining 20% are structured and semi-structured. The produced data are in heterogeneous format and without following any standards. Among heterogeneous (structured, semi-structured and unstructured) data, textual data are nowadays used by industries for prediction and visualization of future challenges. Extracting useful information from it is really challenging for stakeholders due to lexical and semantic matching. Few studies have been solving this issue by using ontologies and semantic tools, but the main limitations of proposed work were the less coverage of multidimensional terms. To solve this problem, this study aims to produce a novel multidimensional reference model using linguistics categories for heterogeneous textual datasets. The categories such context, semantic and syntactic clues are focused along with their score. The main contribution of MRM is that it checks each tokens with each term based on indexing of linguistic categories such as synonym, antonym, formal, lexical word order and co-occurrence. The experiments show that the percentage of MRM is better than the state-of-the-art single dimension reference model in terms of more coverage, linguistics categories and heterogeneous datasets.
翻译:随着技术的发展及最新设备的应用,生产出的数据量巨大。其中,80%的数据是非结构化的,其余20%为结构化和半结构化数据。这些生成的数据格式各异,且未遵循任何标准。在异质(结构化、半结构化和非结构化)数据中,文本数据目前被工业界用于预测和可视化未来挑战。由于词汇和语义匹配问题,利益相关者从中提取有用信息极具挑战性。已有研究通过本体论和语义工具解决此问题,但所提工作的主要局限性在于对多维术语的覆盖不足。为解决此问题,本研究旨在利用语言类别为异质文本数据集提出一种新颖的多维参考模型。重点聚焦于上下文、语义和句法线索等类别及其得分。MRM的主要贡献在于,它基于同义词、反义词、形式词、词汇词序和共现等语言类别的索引,检查每个标记与每个术语的匹配情况。实验表明,在覆盖范围、语言类别和异质数据集方面,MRM的百分比优于当前最先进的单维参考模型。