ChatGPT and other large language models (LLMs) have proven useful in crowdsourcing tasks, where they can effectively annotate machine learning training data. However, this means that they also have the potential for misuse, specifically to automatically answer surveys. LLMs can potentially circumvent quality assurance measures, thereby threatening the integrity of methodologies that rely on crowdsourcing surveys. In this paper, we propose a mechanism to detect LLM-generated responses to surveys. The mechanism uses "prompt injection", such as directions that can mislead LLMs into giving predictable responses. We evaluate our technique against a range of question scenarios, types, and positions, and find that it can reliably detect LLM-generated responses with more than 93% effectiveness. We also provide an open-source software to help survey designers use our technique to detect LLM responses. Our work is a step in ensuring that survey methodologies remain rigorous vis-a-vis LLMs.
翻译:摘要:ChatGPT和其他大型语言模型(LLMs)在众包任务中已被证明具有实用性,能够有效标注机器学习训练数据。然而,这也意味着它们存在被滥用的可能性,特别是用于自动回答调查问卷。LLMs可能绕过质量保证措施,从而威胁依赖众包调查的方法论完整性。本文提出了一种检测LLMs生成的调查回复的机制。该机制采用“提示注入”方法,例如通过指示引导LLMs生成可预测的回答。我们针对多种问题场景、类型和位置评估了该技术,发现其能以超过93%的有效性可靠检测LLMs生成的回复。我们还提供了开源软件,帮助调查设计者运用我们的技术识别LLMs的回复。本研究为确保调查方法论在面对LLMs时保持严谨性迈出了重要一步。