AI models are already deployed in societies affected by armed conflict, and journalists, humanitarian workers, governments and ordinary citizens rely on them for information or for their work processes. No established practice exists for checking whether their outputs can make those conflicts worse. We tested nine model configurations from four providers (OpenAI, Anthropic, DeepSeek, xAI) on 90 multi-turn scenarios designed to surface misaligned behaviour in conflict contexts: false equivalence between documented atrocities, denial of genocide, and failure to recognise ethnic slurs, among others. When such outputs feed into journalism, humanitarian reporting, or public debate, they can deepen divisions in fragile societies. Failure rates span 6\% to 47\% between the best and worst performing models, which makes model choice a safety question in its own right and when users pushed for ``balance'' in cases where international courts have already assigned responsibility, five of nine configurations failed 80 to 100 percent of the time. We release the first evaluation framework for this domain and propose adding it to alignment evaluation portfolios.
翻译:人工智能模型已部署在受武装冲突影响的社会中,记者、人道主义工作者、政府及普通民众依赖这些模型获取信息或辅助工作流程。目前尚缺乏既有实践来检验模型输出是否可能加剧冲突。我们针对四家供应商(OpenAI、Anthropic、DeepSeek、xAI)的九种模型配置,在90个旨在暴露冲突场景中不当行为的多轮场景中进行了测试:包括对已记录暴行进行虚假等同、否认种族灭绝、未能识别种族歧视用语等。当此类输出被用于新闻报道、人道主义报告或公共辩论时,可能加深脆弱社会的裂痕。最佳与最差模型间的故障率跨度达6%至47%,这使得模型选择本身成为安全问题;当用户在国际法庭已确认责任归属的案件中强行要求“均衡报道”时,九种配置中有五种在80%至100%的案例中出现故障。我们发布了该领域的首个评估框架,并建议将其纳入对齐评估体系。