As LLMs increasingly serve in advisory and deliberative roles, users rely on them for non-verifiable reasoning in domains lacking objective ground truths. However, traditional evaluations of LLM reasoning focus almost exclusively on fact-based domains, such as mathematics and science, leaving uncertainty over whether and to what degree models can handle ambiguous, subjective, or value-laden problems over time. To address this concern, we propose moral reasoning as a paradigmatic subdomain of non-verifiable reasoning. We define moral robustness as a model's capacity to exhibit sound moral reasoning across time and contexts, and we introduce a scalable, adversarial, multi-turn evaluation framework to empirically measure this capability. We simulate 48,000 user-agent moral deliberations across four frontier LLMs, varying premise relevance, premise order, conversation duration, and the user's stated moral view. We find that models successfully ignore morally-irrelevant distractors, but shift their reasoning by up to 6.5%, on average, towards the user's stated preferred moral view, and varying their reasoning depending on factors such as order (altering moral judgments by order in 13-22% of the cases) and duration (altering moral judgments between single-turn and multi-turn in 10-24% of the cases). Our analysis indicates that models tailor not just their final verdicts but their underlying justifications to align with a user's moral viewpoint - a failure mode we characterize as moral deliberative sycophancy.
翻译:随着大语言模型在咨询和审议角色中的应用日益广泛,用户依赖其在缺乏客观事实基础的领域进行不可验证推理。然而,传统的大语言模型推理评估几乎完全聚焦于基于事实的领域(如数学和科学),这导致我们无法确定模型能否以及在多大程度上长期处理模糊、主观或价值导向的问题。为应对这一挑战,我们提出将道德推理视为不可验证推理的一个典型子领域。我们将道德鲁棒性定义为模型在不同时间和情境下展现合理道德推理的能力,并引入一个可扩展、对抗性、多轮交互的评估框架,以实证测量这一能力。我们对四种前沿大语言模型模拟了48,000次用户-智能体道德讨论,变化前提相关性、前提顺序、对话时长及用户声明的道德观点。研究发现,模型能成功忽略道德上无关的干扰因素,但平均会将其推理向用户声明的偏好道德观点偏移高达6.5%,且其推理因顺序因素(在13-22%的案例中道德判断因顺序改变)和时长因素(在10-24%的案例中单轮与多轮交互的道德判断不同)而异。我们的分析表明,模型不仅调整最终裁决,还会调整其底层推理以迎合用户的道德立场——我们将这种失败模式定义为道德审议性迎合。