With the rising popularity of user-generated genealogical family trees, new genealogical information systems have been developed. State-of-the-art natural question answering algorithms use deep neural network (DNN) architecture based on self-attention networks. However, some of these models use sequence-based inputs and are not suitable to work with graph-based structure, while graph-based DNN models rely on high levels of comprehensiveness of knowledge graphs that is nonexistent in the genealogical domain. Moreover, these supervised DNN models require training datasets that are absent in the genealogical domain. This study proposes an end-to-end approach for question answering using genealogical family trees by: 1) representing genealogical data as knowledge graphs, 2) converting them to texts, 3) combining them with unstructured texts, and 4) training a trans-former-based question answering model. To evaluate the need for a dedicated approach, a comparison between the fine-tuned model (Uncle-BERT) trained on the auto-generated genealogical dataset and state-of-the-art question-answering models was per-formed. The findings indicate that there are significant differences between answering genealogical questions and open-domain questions. Moreover, the proposed methodology reduces complexity while increasing accuracy and may have practical implications for genealogical research and real-world projects, making genealogical data accessible to experts as well as the general public.
翻译:随着用户生成家谱族谱的日益普及,新型家谱信息系统相继问世。当前最先进的自然语言问答算法采用基于自注意力网络的深度神经网络架构。然而,部分模型依赖序列化输入而无法处理图结构数据,基于图结构的深度神经网络模型又要求知识图谱具有高完备性——这在谱系领域难以实现。此外,这类监督式深度神经网络模型需要训练数据集,而谱系领域恰恰缺乏此类标注数据。本研究提出了一种端到端的家谱族谱问答方法,具体包括:1) 将谱系数据表征为知识图谱,2) 将其转化为文本形式,3) 与非结构化文本进行融合,4) 训练基于Transformer架构的问答模型。为验证专用方法的必要性,我们将在自动生成的谱系数据集上微调的Uncle-BERT模型与当前主流问答模型进行对比实验。研究结果表明,谱系问答与开放域问答存在显著差异。该方法在提升准确率的同时降低了模型复杂度,对谱系研究及实际工程项目具有实践价值,能够使专家及普通大众均可便捷获取谱系数据。