Large language models (LLMs) have significantly advanced the field of natural language processing (NLP), providing a highly useful, task-agnostic foundation for a wide range of applications. The great promise of LLMs as general task solvers motivated people to extend their functionality largely beyond just a ``chatbot'', and use it as an assistant or even replacement for domain experts and tools in specific domains such as healthcare, finance, and education. However, directly applying LLMs to solve sophisticated problems in specific domains meets many hurdles, caused by the heterogeneity of domain data, the sophistication of domain knowledge, the uniqueness of domain objectives, and the diversity of the constraints (e.g., various social norms, cultural conformity, religious beliefs, and ethical standards in the domain applications). To fill such a gap, explosively-increase research, and practices have been conducted in very recent years on the domain specialization of LLMs, which, however, calls for a comprehensive and systematic review to better summarizes and guide this promising domain. In this survey paper, first, we propose a systematic taxonomy that categorizes the LLM domain-specialization techniques based on the accessibility to LLMs and summarizes the framework for all the subcategories as well as their relations and differences to each other. We also present a comprehensive taxonomy of critical application domains that can benefit from specialized LLMs, discussing their practical significance and open challenges. Furthermore, we offer insights into the current research status and future trends in this area.
翻译:大型语言模型(LLM)显著推进了自然语言处理(NLP)领域的发展,为广泛的应用提供了高度实用且任务无关的基础。LLM作为通用任务求解器的巨大潜力,促使人们将其功能大幅扩展至"聊天机器人"之外,并用作特定领域(如医疗、金融、教育)中领域专家和工具的辅助甚至替代。然而,直接应用LLM解决特定领域的复杂问题面临诸多障碍,这些障碍源于领域数据的异质性、领域知识的复杂性、领域目标的独特性以及约束条件的多样性(例如领域应用中各种社会规范、文化契合、宗教信仰和伦理标准)。为弥合这一差距,近年针对LLM领域专精化的研究与实践呈爆炸式增长,但这亟需全面系统的综述来更好地总结和指导这一前景广阔的研究领域。本综述首先提出一个系统化分类法,基于对LLM的可访问性对领域专精化技术进行分类,并总结所有子类别的框架及其相互关系与差异。我们还提出一个关键应用领域的全面分类,这些领域可从专精化LLM中获益,并探讨其实际意义与开放挑战。最后,我们对该领域的当前研究状况和未来趋势提出见解。