The scientific claims drawn from LLM social simulations should be no stronger than the robustness audits that support them. Generative agents bring new expressive power to agent-based modeling, enabling simulations of collective social processes like cooperation, polarization, and norm formation. Yet they also introduce complexity through additional architectural choices, such as agent specification, memory representation, interaction protocols, and environment design. Small perturbations that appear minor to researchers can cascade into macro-level outcomes through repeated interaction, creating a "butterfly effect." Consequently, scientific claims drawn from LLM social simulations may reflect implementation artifacts rather than the social mechanisms being modeled. We support this position with two case studies: a repeated Prisoner's Dilemma and a social media echo chamber simulation. Across multiple models, minor perturbations in persona format and game-instruction framing shift cooperation rates by up to 76 percentage points, while network homophily and hub assignment produce significant and consistent shifts in polarization metrics. We also find that sensitivity is unevenly distributed across both architectural choices and model families: the same perturbation that produces the 76 pp shift in one frontier model only shifts another by 1 pp. Robustness is therefore a property that should be measured per claim and per model, not assumed. To address this validation gap, we introduce TRAILS (Taxonomy for Robustness Audits In LLM Simulations), a robustness-audit taxonomy spanning three levels of simulation design: agent (micro-level), interaction (meso-level), and system (macro-level). We call for robustness to become a first-order validation requirement before LLM social simulations are used to explain mechanisms, evaluate interventions, or inform decisions.
翻译:从大语言模型(LLM)社会模拟中得出的科学论断,其强度不应超过支撑它们的鲁棒性审计。生成式智能体为基于智能体的建模带来了新的表达能力,使得模拟合作、两极分化和规范形成等集体社会过程成为可能。然而,它们也通过额外的架构选择(例如智能体规范、记忆表征、交互协议和环境设计)引入了复杂性。对研究者而言看似微小的扰动,可能通过重复交互级联放大为宏观层面的结果,从而产生“蝴蝶效应”。因此,从LLM社会模拟中得出的科学论断可能反映的是实现伪影,而非被建模的社会机制。我们通过两个案例研究来支持这一观点:一个重复囚徒困境模拟和一个社交媒体回音室模拟。在多个模型中,人物设定格式和游戏指令框架的微小扰动使合作率的变化高达76个百分点,而网络同质性和枢纽节点分配则导致极化指标出现显著且一致的偏移。我们还发现,敏感性在架构选择和模型家族之间分布不均:导致前沿模型中产生76个百分点偏移的同一扰动,在另一模型中仅产生1个百分点的偏移。因此,鲁棒性应针对每个论断和每个模型进行测量,而非想当然。为弥补这一验证缺口,我们引入了TRAILS(大语言模型模拟中鲁棒性审计的分类体系),这是一个涵盖模拟设计三个层次——智能体(微观层)、交互(中观层)和系统(宏观层)——的鲁棒性审计分类体系。我们呼吁,在利用LLM社会模拟来解释机制、评估干预措施或为决策提供依据之前,鲁棒性应成为首要的验证要求。