Natural language understanding (NLU) is integral to task-oriented dialog systems, but demands a considerable amount of annotated training data to increase the coverage of diverse utterances. In this study, we report the construction of a linguistic resource named FIAD (Financial Annotated Dataset) and its use to generate a Korean annotated training data for NLU in the banking customer service (CS) domain. By an empirical examination of a corpus of banking app reviews, we identified three linguistic patterns occurring in Korean request utterances: TOPIC (ENTITY, FEATURE), EVENT, and DISCOURSE MARKER. We represented them in LGGs (Local Grammar Graphs) to generate annotated data covering diverse intents and entities. To assess the practicality of the resource, we evaluate the performances of DIET-only (Intent: 0.91 /Topic [entity+feature]: 0.83), DIET+ HANBERT (I:0.94/T:0.85), DIET+ KoBERT (I:0.94/T:0.86), and DIET+ KorBERT (I:0.95/T:0.84) models trained on FIAD-generated data to extract various types of semantic items.
翻译:自然语言理解(NLU)是任务导向型对话系统的核心组成部分,但需要大量标注训练数据以提升对多样化话语的覆盖度。本研究报告了名为FIAD(金融标注数据集)的语言资源构建过程,及其在银行客服(CS)领域韩语NLU标注训练数据生成中的应用。通过对银行应用评论语料库的实证分析,我们识别出韩语请求语句中出现的三种语言模式:主题(实体、特征)、事件和话语标记。我们将这些模式用局部语法图(LGGs)表示,以生成覆盖多样意图和实体的标注数据。为评估该资源的实用性,我们分别测试了基于FIAD生成数据训练的DIET-only(意图:0.91 /主题[实体+特征]:0.83)、DIET+HANBERT(意图:0.94/主题:0.85)、DIET+KoBERT(意图:0.94/主题:0.86)及DIET+KorBERT(意图:0.95/主题:0.84)模型在提取各类语义项时的性能表现。