We investigated the potential of large language models (LLMs) in developing dataset validation tests. We carried out 96 experiments each for both GPT-3.5 and GPT-4, examining different prompt scenarios, learning modes, temperature settings, and roles. The prompt scenarios were: 1) Asking for expectations, 2) Asking for expectations with a given context, 3) Asking for expectations after requesting a simulation, and 4) Asking for expectations with a provided data sample. For learning modes, we tested: 1) zero-shot, 2) one-shot, and 3) few-shot learning. We also tested four temperature settings: 0, 0.4, 0.6, and 1. Furthermore, two distinct roles were considered: 1) "helpful assistant", 2) "expert data scientist". To gauge consistency, every setup was tested five times. The LLM-generated responses were benchmarked against a gold standard suite, created by an experienced data scientist knowledgeable about the data in question. We find there are considerable returns to the use of few-shot learning, and that the more explicit the data setting can be the better. The best LLM configurations complement, rather than substitute, the gold standard results. This study underscores the value LLMs can bring to the data cleaning and preparation stages of the data science workflow.
翻译:我们探究了大语言模型(LLM)在开发数据集验证测试中的潜力。针对GPT-3.5和GPT-4各进行了96次实验,考察了不同提示场景、学习模式、温度参数及角色设置。提示场景包括:1) 仅要求预期结果,2) 在给定上下文下要求预期结果,3) 先请求模拟再要求预期结果,4) 提供数据样本后要求预期结果。学习模式测试了:1) 零样本学习,2) 单样本学习,3) 少样本学习。同时测试了四个温度参数:0、0.4、0.6和1。此外,对比了两种不同角色:1) "有益助手",2) "数据科学专家"。为评估一致性,每种实验设置均重复五次。将LLM生成的结果与由熟悉相关数据的资深数据科学家构建的黄金标准套件进行对比。研究发现少样本学习能带来显著提升,且数据场景越明确效果越好。最优LLM配置可补充而非替代黄金标准结果。本研究凸显了LLM在数据科学工作流的数据清洗与准备阶段的重要价值。