Improving policymaking is a central concern in public administration. Prior human subject studies reveal substantial cross-cultural differences in citizens' emotional responses to red tape during policy implementation. While LLM agents offer opportunities to simulate human-like responses and reduce experimental costs, their ability to generate culturally appropriate emotional responses to red tape remains unverified. To address this gap, we propose an evaluation framework for assessing LLMs' emotional responses to red tape across diverse cultural contexts. As a pilot study, we apply this framework to a single red-tape scenario. Our results show that all models exhibit limited alignment with human emotional responses, with notably weaker performance in Eastern cultures. Cultural prompting strategies prove largely ineffective in improving alignment. We further introduce \textbf{RAMO}, an interactive interface for simulating citizens' emotional responses to red tape and for collecting human data to improve models. The interface is publicly available at https://ramo-chi.ivia.ch.
翻译:改进政策制定是公共管理的核心关切。先前的人类主体研究表明,在政策实施过程中,不同文化背景下公民对官僚红头文件的情绪反应存在显著差异。虽然大型语言模型代理有望模拟人类反应并降低实验成本,但其生成符合文化情境的官僚红头文件情绪反应的能力尚未得到验证。为填补这一空白,我们提出了一套评估框架,用于评估LLM在不同文化背景下对官僚红头文件的情绪反应。作为试点研究,我们将该框架应用于单一红头文件场景。结果显示,所有模型与人类情绪反应的契合度有限,在东方文化中的表现尤为薄弱。文化提示策略在提升契合度方面基本无效。我们进一步介绍了\textbf{RAMO}——一个用于模拟公民对官僚红头文件的情绪反应以及收集人类数据以改进模型的交互式界面。该界面公开访问地址为https://ramo-chi.ivia.ch。