Model attribution for machine-generated disinformation poses a significant challenge in understanding its origins and mitigating its spread. This task is especially challenging because modern large language models (LLMs) produce disinformation with human-like quality. Additionally, the diversity in prompting methods used to generate disinformation complicates accurate source attribution. These methods introduce domain-specific features that can mask the fundamental characteristics of the models. In this paper, we introduce the concept of model attribution as a domain generalization problem, where each prompting method represents a unique domain. We argue that an effective attribution model must be invariant to these domain-specific features. It should also be proficient in identifying the originating models across all scenarios, reflecting real-world detection challenges. To address this, we introduce a novel approach based on Supervised Contrastive Learning. This method is designed to enhance the model's robustness to variations in prompts and focuses on distinguishing between different source LLMs. We evaluate our model through rigorous experiments involving three common prompting methods: ``open-ended'', ``rewriting'', and ``paraphrasing'', and three advanced LLMs: ``llama 2'', ``chatgpt'', and ``vicuna''. Our results demonstrate the effectiveness of our approach in model attribution tasks, achieving state-of-the-art performance across diverse and unseen datasets.
翻译:机器生成虚假信息的模型溯源对于理解其来源和遏制其传播构成了重大挑战。由于现代大语言模型(LLMs)能够生成具有类人质量的虚假信息,这项任务尤为困难。此外,用于生成虚假信息的提示方法多样性使得准确溯源变得复杂。这些方法引入了领域特定的特征,可能掩盖模型的基本特性。在本文中,我们将模型溯源问题定义为领域泛化问题,其中每种提示方法代表一个独特的领域。我们认为,一个有效的溯源模型必须对这些领域特定特征保持不变性,同时应擅长在所有场景中识别来源模型,以反映现实世界中的检测挑战。为此,我们提出了一种基于监督对比学习的新方法。该方法旨在增强模型对提示变化的鲁棒性,并专注于区分不同的源LLMs。我们通过严谨的实验评估了我们的模型,实验涉及三种常见提示方法:"开放式生成"、"重写"和"释义",以及三种先进LLMs:"llama 2"、"chatgpt"和"vicuna"。我们的结果表明,该方法在模型溯源任务中具有显著效果,在多样化和未见数据集上均达到了最先进的性能。