The first step towards digitalization within organizations lies in digitization - the conversion of analog data into digitally stored data. This basic step is the prerequisite for all following activities like the digitalization of processes or the servitization of products or offerings. However, digitization itself often leads to 'data-rich' but 'knowledge-poor' material. Knowledge discovery and knowledge extraction as approaches try to increase the usefulness of digitized data. In this paper, we point out the key challenges in the context of knowledge discovery and present an approach to addressing these using a microservices architecture. Our solution led to a conceptual design focusing on keyword extraction, similarity calculation of documents, database queries in natural language, and programming language independent provision of the extracted information. In addition, the conceptual design provides referential design guidelines for integrating processes and applications for semi-automatic learning, editing, and visualization of ontologies. The concept also uses a microservices architecture to address non-functional requirements, such as scalability and resilience. The evaluation of the specified requirements is performed using a demonstrator that implements the concept. Furthermore, this modern approach is used in the German patent office in an extended version.
翻译:组织内部数字化的第一步在于数字化——将模拟数据转换为数字存储数据。这一基础步骤是所有后续活动(如流程数字化或产品/服务化)的先决条件。然而,数字化本身往往导致"数据丰富"但"知识贫乏"的材料。知识发现与知识提取作为方法论,旨在提升数字化数据的可用性。本文指出了知识发现背景下的关键挑战,并提出了一种采用微服务架构应对这些挑战的方法。我们的解决方案形成了聚焦于关键词提取、文档相似度计算、自然语言数据库查询以及编程语言无关的信息提取供给的概念设计。此外,该概念设计为半自动本体学习、编辑与可视化的流程及应用集成提供了参考性设计指南。该架构同时采用微服务模式以满足非功能性需求,如可扩展性与弹性。通过实现该概念的演示系统对既定需求进行评估验证。该现代方法已在德国专利局以扩展版本投入实际应用。