The emergence of Large Language Models (LLMs) has great potential to reshape the landscape of many social media platforms. While this can bring promising opportunities, it also raises many threats, such as biases and privacy concerns, and may contribute to the spread of propaganda by malicious actors. We developed the "LLMs Among Us" experimental framework on top of the Mastodon social media platform for bot and human participants to communicate without knowing the ratio or nature of bot and human participants. We built 10 personas with three different LLMs, GPT-4, LLama 2 Chat, and Claude. We conducted three rounds of the experiment and surveyed participants after each round to measure the ability of LLMs to pose as human participants without human detection. We found that participants correctly identified the nature of other users in the experiment only 42% of the time despite knowing the presence of both bots and humans. We also found that the choice of persona had substantially more impact on human perception than the choice of mainstream LLMs.
翻译:大语言模型(LLMs)的出现具有重塑众多社交媒体平台格局的巨大潜力。虽然这能够带来令人期待的机会,但也引发了许多威胁,例如偏见和隐私问题,并可能助长恶意行为者传播宣传内容。我们在Mastodon社交媒体平台上构建了“我们之中的LLMs”实验框架,使机器人和人类参与者能够在不了解彼此比例或性质的情况下进行交流。我们利用三种不同的大语言模型——GPT-4、LLama 2 Chat和Claude——创建了10种人物角色。我们进行了三轮实验,并在每轮实验后对参与者进行调查,以衡量大语言模型在未被人为检测的情况下伪装成人类参与者的能力。结果发现,尽管参与者知道实验中有机器人和人类同时存在,他们正确识别其他用户性质的准确率仅为42%。我们还发现,人物角色的选择对人类感知的影响远大于主流大语言模型的选择。