Traditional dataset retrieval systems index on metadata information rather than on the data values. Thus relying primarily on manual annotations and high-quality metadata, processes known to be labour-intensive and challenging to automate. We propose a method to support metadata enrichment with topic annotations of column headers using three Large Language Models (LLMs): ChatGPT-3.5, GoogleBard and GoogleGemini. We investigate the LLMs ability to classify column headers based on domain-specific topics from a controlled vocabulary. We evaluate our approach by assessing the internal consistency of the LLMs, the inter-machine alignment, and the human-machine agreement for the topic classification task. Additionally, we investigate the impact of contextual information (i.e. dataset description) on the classification outcomes. Our results suggest that ChatGPT and GoogleGemini outperform GoogleBard for internal consistency as well as LLM-human-alignment. Interestingly, we found that context had no impact on the LLMs performances. This work proposes a novel approach that leverages LLMs for text classification using a controlled topic vocabulary, which has the potential to facilitate automated metadata enrichment, thereby enhancing dataset retrieval and the Findability, Accessibility, Interoperability and Reusability (FAIR) of research data on the Web.
翻译:传统数据集检索系统基于元数据信息而非数据值进行索引,因此主要依赖人工标注和高质量元数据——这些过程以劳动密集且难以自动化为特征。我们提出一种方法,通过使用三种大型语言模型(ChatGPT-3.5、GoogleBard和GoogleGemini),利用列标题的主题标注来支持元数据增强。我们研究了这些模型基于受控词表中特定领域主题对列标题进行分类的能力。通过评估LLMs的内部一致性、跨模型对齐性以及人机一致性来评价分类任务效果,同时探究上下文信息(如数据集描述)对分类结果的影响。实验结果表明,ChatGPT和GoogleGemini在内部一致性及人机对齐方面均优于GoogleBard。值得注意的是,我们发现上下文信息对LLMs的性能无显著影响。本研究提出了一种利用受控主题词表进行文本分类的创新方法,该方法有望实现元数据自动增强,从而提升数据集检索效率以及网络研究数据的可发现性、可访问性、互操作性和可重复使用性(FAIR原则)。