Simulations, and more recently LLM agent simulations, have been adopted as useful tools for policymakers to explore interventions, rehearse potential scenarios, and forecast outcomes. While LLM simulations have enormous potential, two critical challenges remain understudied: the dual-use potential of accurate models of individual or population-level human behavior and the difficulty of validating simulation outputs. In light of these limitations, we must define boundaries for both simulation developers and decision-makers to ensure responsible development and ethical use. We propose and discuss three preconditions for societal-scale LLM agent simulations: 1) do not treat simulations of marginalized populations as neutral technical outputs, 2) do not simulate populations without their participation, and 3) do not simulate without accountability. We believe that these guardrails, combined with our call for simulation development and deployment reports, will help build trust among policymakers while promoting responsible development and use of societal-scale LLM agent simulations for the public benefit.
翻译:模拟(尤其是近期的大语言模型智能体模拟)已被政策制定者广泛采用,作为探索干预措施、预演潜在场景及预测结果的有效工具。尽管大语言模型模拟具有巨大潜力,但两个关键挑战仍未得到充分研究:个体或群体层面人类行为精准模型的双重用途可能性,以及模拟输出结果的验证困难。鉴于这些局限性,我们必须为模拟开发者和决策者界定边界,以确保负责任的开发与合乎伦理的使用。我们提出并讨论了社会规模大语言模型智能体模拟的三项前提条件:1)不得将边缘群体的模拟视为中立技术产出;2)不得在未获得群体参与的情况下进行模拟;3)不得在缺乏问责机制的情况下进行模拟。我们相信,这些防护措施结合对模拟开发与部署报告的倡议,将有助于在政策制定者间建立信任,同时促进以公共利益为导向的社会规模大语言模型智能体模拟的负责任开发与使用。