This paper presents an overview of the PromptCBLUE shared task (http://cips-chip.org.cn/2023/eval1) held in the CHIP-2023 Conference. This shared task reformualtes the CBLUE benchmark, and provide a good testbed for Chinese open-domain or medical-domain large language models (LLMs) in general medical natural language processing. Two different tracks are held: (a) prompt tuning track, investigating the multitask prompt tuning of LLMs, (b) probing the in-context learning capabilities of open-sourced LLMs. Many teams from both the industry and academia participated in the shared tasks, and the top teams achieved amazing test results. This paper describes the tasks, the datasets, evaluation metrics, and the top systems for both tasks. Finally, the paper summarizes the techniques and results of the evaluation of the various approaches explored by the participating teams.
翻译:本文介绍了在 CHIP-2023 会议上举办的 PromptCBLUE 共享任务(http://cips-chip.org.cn/2023/eval1)。该共享任务重新设计了 CBLUE 基准测试,为面向通用医学自然语言处理的中文开放域或医学领域大语言模型(LLMs)提供了良好的测试平台。任务设有两个不同赛道:(a)提示微调赛道,研究基于多任务提示的 LLMs 微调方法;(b)探索开源 LLMs 的上下文学习能力。来自工业界和学术界的众多团队参与了该共享任务,其中表现最佳的团队取得了令人瞩目的测试结果。本文详细描述了这两个赛道的任务设置、数据集、评估指标以及最优系统方案。最后,文章总结了各参赛团队所探索的多种技术路线及其评估结果。