Investigating Public Fine-Tuning Datasets: A Complex Review of Current Practices from a Construction Perspective

With the rapid development of the large model domain, research related to fine-tuning has concurrently seen significant advancement, given that fine-tuning is a constituent part of the training process for large-scale models. Data engineering plays a fundamental role in the training process of models, which includes data infrastructure, data processing, etc. Data during fine-tuning likewise forms the base for large models. In order to embrace the power and explore new possibilities of fine-tuning datasets, this paper reviews current public fine-tuning datasets from the perspective of data construction. An overview of public fine-tuning datasets from two sides: evolution and taxonomy, is provided in this review, aiming to chart the development trajectory. Construction techniques and methods for public fine-tuning datasets of Large Language Models (LLMs), including data generation and data augmentation among others, are detailed. This elaboration follows the aforementioned taxonomy, specifically across demonstration, comparison, and generalist categories. Additionally, a category tree of data generation techniques has been abstracted in our review to assist researchers in gaining a deeper understanding of fine-tuning datasets from the construction dimension. Our review also summarizes the construction features in different data preparation phases of current practices in this field, aiming to provide a comprehensive overview and inform future research. Fine-tuning dataset practices, encompassing various data modalities, are also discussed from a construction perspective in our review. Towards the end of the article, we offer insights and considerations regarding the future construction and developments of fine-tuning datasets.

翻译：随着大模型领域的快速发展，作为大规模模型训练过程组成部分的微调相关研究也取得了显著进展。数据工程在模型训练过程中发挥着基础性作用，包括数据基础设施、数据处理等方面。微调阶段的数据同样构成大模型的基础。为充分发挥微调数据集的潜力并探索其新的可能性，本文从数据构建的视角对当前公开微调数据集进行了系统性综述。本综述从演变历程和分类体系两个维度概述公开微调数据集，旨在描绘其发展轨迹。详细阐述了大型语言模型（LLMs）公开微调数据集的构建技术和方法，包括数据生成、数据增强等。这些阐述遵循前述分类体系，具体涵盖示范型、比较型和通用型三大类别。此外，我们在综述中抽象出数据生成技术的分类树，以帮助研究者从构建维度深入理解微调数据集。本综述还总结了该领域当前实践在不同数据准备阶段的构建特征，旨在提供全面概览并为未来研究提供参考。我们从构建视角探讨了涵盖多种数据模态的微调数据集实践案例。在文章结尾部分，我们针对微调数据集未来的构建方向与发展趋势提出了见解与思考。

相关内容

MoDELS

关注 45

ACM/IEEE第23届模型驱动工程语言和系统国际会议，是模型驱动软件和系统工程的首要会议系列，由ACM-SIGSOFT和IEEE-TCSE支持组织。自1998年以来，模型涵盖了建模的各个方面，从语言和方法到工具和应用程序。模特的参加者来自不同的背景，包括研究人员、学者、工程师和工业专业人士。MODELS 2019是一个论坛，参与者可以围绕建模和模型驱动的软件和系统交流前沿研究成果和创新实践经验。今年的版本将为建模社区提供进一步推进建模基础的机会，并在网络物理系统、嵌入式系统、社会技术系统、云计算、大数据、机器学习、安全、开源等新兴领域提出建模的创新应用以及可持续性。官网链接：http://www.modelsconference.org/

【CVPR 2022】一个完全无监督的框架，从噪声和部分测量中学习图像，Robust Equivariant Imaging: a fully unsupervised framework for learning to image

专知会员服务

25+阅读 · 2022年3月3日

UCM《机器学习导论笔记》，80页pdf CSE176 Introduction to Machine Learning

专知会员服务

32+阅读 · 2021年9月29日

【亚马逊-WWW2020】不解析,生成!用于面向任务的语义分析的序列到序列体系结构，Don't Parse, Generate! A Sequence to Sequence Architecture for Task-Oriented Semantic Parsing

专知会员服务

15+阅读 · 2020年2月1日

FlowQA: Grasping Flow in History for Conversational Machine Comprehension

专知会员服务

34+阅读 · 2019年10月18日