Low-quality data can cause downstream problems in high-stakes applications. Data-centric approach emphasizes on improving dataset quality to enhance model performance. High-quality datasets are needed for general-purpose Large Language Models (LLMs) training, as well as for domain-specific models, which are usually small in size as it is costly to engage a large number of domain experts for their creation. Thus, it is vital to ensure high-quality domain-specific training data. In this paper, we propose a framework for enhancing the data quality of original datasets. We applied the proposed framework to four biomedical datasets and showed relative improvement of up to 33%/40% for fine-tuning of retrieval/reader models on the BioASQ dataset when using back translation to enhance the original dataset quality.
翻译:低质量数据可能在高风险应用中引发下游问题。以数据为中心的方法强调通过提升数据集质量来增强模型性能。通用大语言模型(LLMs)的训练以及领域特定模型都需要高质量数据集,而由于聘请大量领域专家创建数据成本高昂,领域特定模型通常规模较小。因此,确保领域特定训练数据的高质量至关重要。本文提出一个框架,用于提升原始数据集的数据质量。我们将该框架应用于四个生物医学数据集,结果表明,在使用回译增强原始数据集质量时,在BioASQ数据集上对检索/阅读器模型进行微调时,相对性能提升最高可达33%/40%。