Extraction of concepts and entities of interest from non-formal texts such as social media posts and informal communication is an important capability for decision support systems in many domains, including healthcare, customer relationship management, and others. Despite the recent advances in training large language models for a variety of natural language processing tasks, the developed models and techniques have mainly focused on formal texts and do not perform as well on colloquial data, which is characterized by a number of distinct challenges. In our research, we focus on the healthcare domain and investigate the problem of symptom recognition from colloquial texts by designing and evaluating several training strategies for BERT-based model fine-tuning. These strategies are distinguished by the choice of the base model, the training corpora, and application of term perturbations in the training data. The best-performing models trained using these strategies outperform the state-of-the-art specialized symptom recognizer by a large margin. Through a series of experiments, we have found specific patterns of model behavior associated with the training strategies we designed. We present design principles for training strategies for effective entity recognition in colloquial texts based on our findings.
翻译:从非正式文本(如社交媒体帖子和非正式交流)中提取感兴趣的概念和实体,是许多领域(包括医疗保健、客户关系管理等)决策支持系统的重要能力。尽管近期在训练大型语言模型以处理各种自然语言处理任务方面取得了进展,但所开发的模型和技术主要针对正式文本,在具有诸多独特挑战的口语数据上表现不佳。在我们的研究中,我们聚焦于医疗保健领域,通过设计和评估几种基于BERT模型微调的训练策略,探讨从口语文本中识别症状的问题。这些策略的区别在于基础模型的选择、训练语料库以及在训练数据中应用术语扰动。使用这些策略训练出的最佳模型,在性能上大幅超越了当前最先进的特异性症状识别器。通过一系列实验,我们发现了与我们设计的训练策略相关的特定模型行为模式。基于我们的发现,我们提出了针对口语文本中有效实体识别的训练策略设计原则。